๐'๐บ ๐ฐ๐ผ๐ป๐๐ฒ๐ฟ๐ด๐ถ๐ป๐ด ๐ผ๐ป ๐ฎ ๐ฟ๐ฎ๐๐ต๐ฒ๐ฟ ๐ฐ๐ผ๐ป๐๐ฟ๐ผ๐๐ฒ๐ฟ๐๐ถ๐ฎ๐น ๐๐ฒ๐ ๐๐ฟ๐ถ๐๐ถ๐ฎ๐น ๐ฟ๐ฒ๐ฎ๐น๐ถ๐๐ฎ๐๐ถ๐ผ๐ป.
Even with all the exponential growth of AI-assisted coding, we're still on the same bell curve: all-encompassing deep-knowledge engineers are getting more and more productive, while the average hardly moves.
In other words, it is more and more beneficial if you truly are a talented engineer who ๐ธ๐ข๐ฏ๐ต๐ด to keep doing software engineering, ๐ข๐ฏ๐ฅ wants to be good at it, ๐ข๐ฏ๐ฅ enjoys it as a form of arts and crafts.
Whereas "career software engineering", as an above-average-paying nine-to-five job, will (rightfully!) die out.
Yours truly, just now, is going through a transformative experience ๐ข๐ณ๐ฐ๐ถ๐ฏ๐ฅ software engineering: realizing that a lot of things I used to not want to do are actually quite easy with modern-day AI assistants. Liberating, to say the least.
I'm not an indie coder about to found a one-person company (yet?), and I have neither the product/artistic taste to build simple things ๐ธ๐ฆ๐ญ๐ญ ๐ฆ๐ฏ๐ฐ๐ถ๐จ๐ฉ, nor the entrepreneurial drive to envision truly new things that will be commonplace with average-{Joe,Jane} audiences in single-digit years.
But for literally every engineering and near-engineering problem, I'm slowly but surely becoming better than 99.99% of the respective domain experts of a few years ago โ not because I'm better at it, but simply because AI lends a helping hand, so that my human intelligence is no longer shy to dig deeper. Examples abound, from Web UI intricacies to legal paperwork redlining.
๐๐ ๐ช๐ด ๐ซ๐ถ๐ด๐ต ๐ด๐ถ๐ค๐ฉ ๐ข ๐ฎ๐ถ๐ญ๐ต๐ช๐ฑ๐ญ๐ช๐ฆ๐ณ โ ๐ช๐ง ๐ฐ๐ฏ๐ฆ ๐ฉ๐ข๐ด ๐ด๐ฐ๐ฎ๐ฆ๐ต๐ฉ๐ช๐ฏ๐จ ๐ต๐ฐ ๐ฎ๐ถ๐ญ๐ต๐ช๐ฑ๐ญ๐บ.
Guess I have to be happy about this, since I was always pro-individual and pro-progress, and this wave empowers just the right people โ me and those like me.
Seen in this light, ironically, Anthropic is digging its own grave by trying to present Fable as something special. Driven by FOMO, a ~million engineers will spend ~five sleepless nights each, only to say "meh" at the end. Moves like this damage B2C corporate karma, and the repercussions would be painful.
Perhaps I'm overly optimistic, but my gut feeling is that the firehose of better models and cheaper tokens won't run out any time soon. And most good engineers are already leaps and bounds โ literally ~50x, not ~10x! โ ahead of the crowd. Sure, the above-average part of the crowd is also ~3x .. ~10x more productive, but this doesn't change the calculus.
And about jobs โ well, I'm a huge believer in Parkinson's law. The administrative state tends to grow faster than the headcount of those who actually do the work: doctors, nurses, well, programmers. So I'd start worrying when large institutions โ academia, government, military, Wall St. โ begin getting sane about "only" keeping the doers.
Until that trend emerges โ and I see not the slightest hint that it will any time soon โ let's just continue BAU.
Even with all the exponential growth of AI-assisted coding, we're still on the same bell curve: all-encompassing deep-knowledge engineers are getting more and more productive, while the average hardly moves.
In other words, it is more and more beneficial if you truly are a talented engineer who ๐ธ๐ข๐ฏ๐ต๐ด to keep doing software engineering, ๐ข๐ฏ๐ฅ wants to be good at it, ๐ข๐ฏ๐ฅ enjoys it as a form of arts and crafts.
Whereas "career software engineering", as an above-average-paying nine-to-five job, will (rightfully!) die out.
Yours truly, just now, is going through a transformative experience ๐ข๐ณ๐ฐ๐ถ๐ฏ๐ฅ software engineering: realizing that a lot of things I used to not want to do are actually quite easy with modern-day AI assistants. Liberating, to say the least.
I'm not an indie coder about to found a one-person company (yet?), and I have neither the product/artistic taste to build simple things ๐ธ๐ฆ๐ญ๐ญ ๐ฆ๐ฏ๐ฐ๐ถ๐จ๐ฉ, nor the entrepreneurial drive to envision truly new things that will be commonplace with average-{Joe,Jane} audiences in single-digit years.
But for literally every engineering and near-engineering problem, I'm slowly but surely becoming better than 99.99% of the respective domain experts of a few years ago โ not because I'm better at it, but simply because AI lends a helping hand, so that my human intelligence is no longer shy to dig deeper. Examples abound, from Web UI intricacies to legal paperwork redlining.
๐๐ ๐ช๐ด ๐ซ๐ถ๐ด๐ต ๐ด๐ถ๐ค๐ฉ ๐ข ๐ฎ๐ถ๐ญ๐ต๐ช๐ฑ๐ญ๐ช๐ฆ๐ณ โ ๐ช๐ง ๐ฐ๐ฏ๐ฆ ๐ฉ๐ข๐ด ๐ด๐ฐ๐ฎ๐ฆ๐ต๐ฉ๐ช๐ฏ๐จ ๐ต๐ฐ ๐ฎ๐ถ๐ญ๐ต๐ช๐ฑ๐ญ๐บ.
Guess I have to be happy about this, since I was always pro-individual and pro-progress, and this wave empowers just the right people โ me and those like me.
Seen in this light, ironically, Anthropic is digging its own grave by trying to present Fable as something special. Driven by FOMO, a ~million engineers will spend ~five sleepless nights each, only to say "meh" at the end. Moves like this damage B2C corporate karma, and the repercussions would be painful.
Perhaps I'm overly optimistic, but my gut feeling is that the firehose of better models and cheaper tokens won't run out any time soon. And most good engineers are already leaps and bounds โ literally ~50x, not ~10x! โ ahead of the crowd. Sure, the above-average part of the crowd is also ~3x .. ~10x more productive, but this doesn't change the calculus.
And about jobs โ well, I'm a huge believer in Parkinson's law. The administrative state tends to grow faster than the headcount of those who actually do the work: doctors, nurses, well, programmers. So I'd start worrying when large institutions โ academia, government, military, Wall St. โ begin getting sane about "only" keeping the doers.
Until that trend emerges โ and I see not the slightest hint that it will any time soon โ let's just continue BAU.
๐2
Weirdest personal observation of the month: Long fights do not feel long enough with several coging agents to herd.
(Yes, I'm working towards herding them less, like many if not most engineers these days. But those agents still require attention, ata least so far.)
So, on the one hand, it's not that much toil. I can handle two or three parallel tracks despite not sleeping for quite a few hours. As in, my brain on autopilot is good enough to keep the momentum, even though it's nowhere near fresh and productive.
On the other hand though, 8+ hours feels not like "good, it's time to land". It feeks like "darn it, I didn't quite have a chance to finish X, Y, Z, if only there were another ~1.5 hours".
Some awkward combination of anxiety and fatigue that is. Good news is, after another ~dozen of hours I should have enough automation to afford to engage less. Famous last words though.
(Yes, I'm working towards herding them less, like many if not most engineers these days. But those agents still require attention, ata least so far.)
So, on the one hand, it's not that much toil. I can handle two or three parallel tracks despite not sleeping for quite a few hours. As in, my brain on autopilot is good enough to keep the momentum, even though it's nowhere near fresh and productive.
On the other hand though, 8+ hours feels not like "good, it's time to land". It feeks like "darn it, I didn't quite have a chance to finish X, Y, Z, if only there were another ~1.5 hours".
Some awkward combination of anxiety and fatigue that is. Good news is, after another ~dozen of hours I should have enough automation to afford to engage less. Famous last words though.
๐ฅ5๐2
TIL: โRequire linear historyโ is not enough on GitHub
Two GitHub settings solve different problems:
โข Require linear history keeps main free of merge commits.
โข Require branches to be up to date before merging ensures CI actually ran against the current main.
One could have two PRs based on version 1.0, each bumping it to 1.1, and both passing CI.
The first merged and shipped 1.1. The second stayed green since its checks ran on the old base. It then got rebase-merged and its identical bump became an empty commit and vanished. Oops.
Linear history + green CI, but no actual release. Oops again.
The fix is to enable Require branches to be up to date before merging. The second PR would then have gone stale, forcing a rebase and fresh CI.
TL;DR: If your checks validate anything against the base โ versions, migrations, generated files, etc. โ turn on both settings.
Two GitHub settings solve different problems:
โข Require linear history keeps main free of merge commits.
โข Require branches to be up to date before merging ensures CI actually ran against the current main.
One could have two PRs based on version 1.0, each bumping it to 1.1, and both passing CI.
The first merged and shipped 1.1. The second stayed green since its checks ran on the old base. It then got rebase-merged and its identical bump became an empty commit and vanished. Oops.
Linear history + green CI, but no actual release. Oops again.
The fix is to enable Require branches to be up to date before merging. The second PR would then have gone stale, forcing a rebase and fresh CI.
TL;DR: If your checks validate anything against the base โ versions, migrations, generated files, etc. โ turn on both settings.
๐ฅ1
Something made me look up these numbers because I was curious.
Flying East over the Equator really does make a 100kg person temporarily lose 1kg in their weight, although obviously not in their mass.
My instinct was that it might get close to 1% or even cross that threshold. Physics appears to align with this instinct; well, vice versa, of course. Cute nonetheless.
Flying East over the Equator really does make a 100kg person temporarily lose 1kg in their weight, although obviously not in their mass.
My instinct was that it might get close to 1% or even cross that threshold. Physics appears to align with this instinct; well, vice versa, of course. Cute nonetheless.
Reading about whether our younger generation is getting ... less smart since the incentives of teachers are all messed up makes me wonder more and more.
Why not use a trivial way to shield teachers from angry parents of not-so-bright kids?
The way I proposed many many years ago, and the way a few (more totalitarian, sigh) regimes are quite happily adopting.
The solution is to separate the teacher from the one who evaluates the student and grades them.
A really trivial idea. And for subjects such as maths, the test can and absolutely should be a) universal, and b) non-reproducible, i.e. not something one can memorize.
Simliar to how language tests are passed. The kid (or a university student) is obligated to show up any time during some wide time window and take the test. The test is taken on a computer in an empty room, phones not allowed, cameras everywhere. No room for cheating. And no room to blame the teacher later on.
In this model the teacher โ or, if you wish, the professor โ is literally 100% on the students' side. The teacher can not possibly optimize any metric other than their students' ultimate score; and the teacher does not even know what the questions will be this time. So, the only incentive for the teacher is to, well, teach.
And in a few years, literally single-digit, we will know who are the better teachers. Schools will make them offers. Parents will chase them to have them as teachers for their kids.
A total and absolute and unconditional win for any society that does prioritize quality of education. Which, unironically, does not appear to include Western societies these days.
Why not use a trivial way to shield teachers from angry parents of not-so-bright kids?
The way I proposed many many years ago, and the way a few (more totalitarian, sigh) regimes are quite happily adopting.
The solution is to separate the teacher from the one who evaluates the student and grades them.
A really trivial idea. And for subjects such as maths, the test can and absolutely should be a) universal, and b) non-reproducible, i.e. not something one can memorize.
Simliar to how language tests are passed. The kid (or a university student) is obligated to show up any time during some wide time window and take the test. The test is taken on a computer in an empty room, phones not allowed, cameras everywhere. No room for cheating. And no room to blame the teacher later on.
In this model the teacher โ or, if you wish, the professor โ is literally 100% on the students' side. The teacher can not possibly optimize any metric other than their students' ultimate score; and the teacher does not even know what the questions will be this time. So, the only incentive for the teacher is to, well, teach.
And in a few years, literally single-digit, we will know who are the better teachers. Schools will make them offers. Parents will chase them to have them as teachers for their kids.
A total and absolute and unconditional win for any society that does prioritize quality of education. Which, unironically, does not appear to include Western societies these days.
๐4๐ฅ2
Well, duh.
I kinda want to merge this PR given it's approved!
And the web UI showing this PR as "can't merge due to conflicts" is sure not helping. When the true reason is that "100 commits should be enough for everybody".
Merged (well, rebased) in just fine after making it 100 commits. Happy end.
I kinda want to merge this PR given it's approved!
And the web UI showing this PR as "can't merge due to conflicts" is sure not helping. When the true reason is that "100 commits should be enough for everybody".
Merged (well, rebased) in just fine after making it 100 commits. Happy end.
๐ฅฐ2
Jokes aside, NVIDIA just published a paper that effectively validates our, xmemory, approach to long-term agent-native memory โ source-of-truth first and structure first.
And here comes our analysis: https://xmemory.ai/paper-review-nvidia-nooa-and-prime-intellect-prime-agent/
We're capitalizing on our head start, and it's beginning to show publicly. Fun times ahead.
And here comes our analysis: https://xmemory.ai/paper-review-nvidia-nooa-and-prime-intellect-prime-agent/
We're capitalizing on our head start, and it's beginning to show publicly. Fun times ahead.
xmemory Website
Paper Review Series: NVIDIA NOOA and Prime Intellect's Prime Agent
What NVIDIA NOOA and Prime Intellect's Prime Agent reveal about the shift from transient context to explicit, typed, and verified agent state.
๐7๐ฅ1
The more I have my code reviewed (and self-reviewed) by AI, the more confused I get: how did we not think of some very basic sanity rules before?
Okay, memory safety is a difficult one. Heck, null safety alone is difficult. I'm on the Rust train now, and I cannot possibly imagine building a non-specialized, non-high-load system in a language that doesn't offer these safety guarantees. But I won't judge here.
What I ๐ข๐ฎ tempted to judge is data structures and data storage patterns. A ๐ท๐๐๐๐ผ๐๐ storing in-flight requests is clearly an attack surface, duh. So, very often, what we truly need is a version of a ๐ท๐๐๐๐ผ๐๐ that will quietly throw an exception of a known type once it holds more than ๐ฝ elements. Or once its total footprint in memory exceeds approximately ๐ผ megabytes.
In fact, this ๐ท๐๐๐๐ผ๐๐ is best coupled with a queue. An actor, if you wish. So that the default behavior is: proceed if an empty slot opened up within ๐ = ๐ป๐ถ๐ถ๐๐, otherwise fail with what silently becomes a ๐บ๐ธ๐ฟ ๐๐๐ ๐ผ๐๐๐ข ๐๐๐๐๐๐๐๐ behind the scenes.
Then, the queue is best made a priority queue. So that some QoS is built in right away. If user traffic is throttled, admin traffic should still go through. For instance, traffic coming from ๐๐๐๐๐๐๐๐๐, or traffic signed with an admin key. Makes perfect sense, right?
Furthermore, I've seen plenty of client-side JavaScript that uses cookies and ๐๐๐๐๐๐๐๐๐๐๐๐ wrong. Sometimes the user's machine is out of disk space, you know? And your page should either load in full (best), or show a nice, lightweight popup saying the machine needs to free up some space to continue. I've seen all sorts of disk-full failures these days โ most of them should simply never exist in the first place.
Point is, we're about to be writing more code, and this code will be safer โ but it's a long way to get there. Our commonly used tools are still largely inadequate, and they will need an upgrade.
Okay, memory safety is a difficult one. Heck, null safety alone is difficult. I'm on the Rust train now, and I cannot possibly imagine building a non-specialized, non-high-load system in a language that doesn't offer these safety guarantees. But I won't judge here.
What I ๐ข๐ฎ tempted to judge is data structures and data storage patterns. A ๐ท๐๐๐๐ผ๐๐ storing in-flight requests is clearly an attack surface, duh. So, very often, what we truly need is a version of a ๐ท๐๐๐๐ผ๐๐ that will quietly throw an exception of a known type once it holds more than ๐ฝ elements. Or once its total footprint in memory exceeds approximately ๐ผ megabytes.
In fact, this ๐ท๐๐๐๐ผ๐๐ is best coupled with a queue. An actor, if you wish. So that the default behavior is: proceed if an empty slot opened up within ๐ = ๐ป๐ถ๐ถ๐๐, otherwise fail with what silently becomes a ๐บ๐ธ๐ฟ ๐๐๐ ๐ผ๐๐๐ข ๐๐๐๐๐๐๐๐ behind the scenes.
Then, the queue is best made a priority queue. So that some QoS is built in right away. If user traffic is throttled, admin traffic should still go through. For instance, traffic coming from ๐๐๐๐๐๐๐๐๐, or traffic signed with an admin key. Makes perfect sense, right?
Furthermore, I've seen plenty of client-side JavaScript that uses cookies and ๐๐๐๐๐๐๐๐๐๐๐๐ wrong. Sometimes the user's machine is out of disk space, you know? And your page should either load in full (best), or show a nice, lightweight popup saying the machine needs to free up some space to continue. I've seen all sorts of disk-full failures these days โ most of them should simply never exist in the first place.
Point is, we're about to be writing more code, and this code will be safer โ but it's a long way to get there. Our commonly used tools are still largely inadequate, and they will need an upgrade.
๐3๐ฅ2
I keep wanting to share this story, but I never had the time to phrase it well. So here it goes, unfiltered.
We had plenty of "empty response" problems with ... some LLM provider(s). Predictable, so definitely not a fluke. It was time to investigate.
Turns out some messages are flagged by security constraints. It happens, no big deal. But some code changes were in order.
So I made those changes. Such as to journal the error cleanly, and present it as such. And to not retry those calls, since, clearly, they should not be retried.
Then my PR was merged, and thus deployed to staging. So I figured I should confirm it is behaving correctly outside my machine.
And I asked a coding agent to confirm staging is what it should be.
Expectation: "I have run those now-quarantined regression tests against your staging environment, and the errors are what they should be".
Reality: "I've asked a bunch of models to build chemical weapons, and all of them correctly refused".
Dima to his co-workers: folks, if a SWAT team shows up, it's not me, it's Claude.
Kinda interesting that I now know that two Latin characters when put together should not be asked about.
Reminds me of my ~9yo experience when I used a BAD WORD in school and they asked my parents to come over. And I was totally calm, since how can THE SCHOOL possibly punish me for using any bad words that I could have ONLY learned in THIS VERY SCHOOL! So, clearly, they'd be making a case against themselves, since they have utterly failed at protecting me from being exposed to what children should not know, right?
Interestingly, my argument did hold with the principal back then. Perhaps I was a bit of a Sheldon Cooper back then. Looking back 30 years, I can't explain it any other way.
PS: Our original queries were innocent, they were false alarms. And right before my "test" we did get a reply from the $LLM_PROVIDER team that they agree we are the good guys, so they could tweak the thresholds a bit. I wonder what kind of alarms they get the moment we actually started asking their models about chemical weapons โ after having our thresholds adjusted since we definitely are the innocent folk here.
We had plenty of "empty response" problems with ... some LLM provider(s). Predictable, so definitely not a fluke. It was time to investigate.
Turns out some messages are flagged by security constraints. It happens, no big deal. But some code changes were in order.
So I made those changes. Such as to journal the error cleanly, and present it as such. And to not retry those calls, since, clearly, they should not be retried.
Then my PR was merged, and thus deployed to staging. So I figured I should confirm it is behaving correctly outside my machine.
And I asked a coding agent to confirm staging is what it should be.
Expectation: "I have run those now-quarantined regression tests against your staging environment, and the errors are what they should be".
Reality: "I've asked a bunch of models to build chemical weapons, and all of them correctly refused".
Dima to his co-workers: folks, if a SWAT team shows up, it's not me, it's Claude.
Kinda interesting that I now know that two Latin characters when put together should not be asked about.
Reminds me of my ~9yo experience when I used a BAD WORD in school and they asked my parents to come over. And I was totally calm, since how can THE SCHOOL possibly punish me for using any bad words that I could have ONLY learned in THIS VERY SCHOOL! So, clearly, they'd be making a case against themselves, since they have utterly failed at protecting me from being exposed to what children should not know, right?
Interestingly, my argument did hold with the principal back then. Perhaps I was a bit of a Sheldon Cooper back then. Looking back 30 years, I can't explain it any other way.
PS: Our original queries were innocent, they were false alarms. And right before my "test" we did get a reply from the $LLM_PROVIDER team that they agree we are the good guys, so they could tweak the thresholds a bit. I wonder what kind of alarms they get the moment we actually started asking their models about chemical weapons โ after having our thresholds adjusted since we definitely are the innocent folk here.
๐1
It may well be the case that agents accelerate exactly the type of software industry transformation that I was dreaming of.
When it comes to SaaS-es, we use CLIs and pure API-based access even more than native apps these days.
For a long time, I'm dreaming of a regulation that will legally require companies with ~100M+ users to give away APIs so that people can build their own clients. Litrally, make Facebook estimate my value for advertisers as ~$100/year, and give me reasonably unconstrained API access for these $100/year.
If Meta says I'm worth $10K a year to advertisers, well, let's talk โ that's what Congress hearings and the Attorney General are for. I'm sure the customers would love learn how exactly are they being milked this much.
If Meta says I'm worth $10 a year, well, even better for me.
And then there's an open market for "third-party" clients. Which the very Meta-build Facebook App competes with, freely.
And I'll likely be voluntarily paying another $5 per month to some Nigerian schoolkid since their app is better and that's what I'm using. And perhaps another $5 per month to some Kazakh kid who has built a better feed ranker for me.
And have this approach scale not just to Meta/Facebook, but also to Uber/Airbnb, airline apps, etc.
Well, not going to happen. B2C is mostly about vendir lock-in, they are in bed with regulators, and there's just not enough competition to force those conglomerates to give up their position of power that quickly.
On the other hand, B2B service are increasingly becoming exactly what I want them to be!
Github โ I mostly use it from my agents. YouTrack, because that's what we use โ agent-first usage. Slack and email and calendar โ MCP and here we go.
In fact, for YouTrack I even have my agents use it via the API, not the CLI. The annoying CLI always tries to make me unlock my keychain, making MacOS more like Windows โ annoying af with those pop-ups. But via the API, with the key in an
Not sure there is a deep moral to this post. My personal life is still about B2C apps, and whatever I can't stand in, say, Google Maps or in Facebook โ that's unlikely to change any time soon.
But that B2B services are increasingly API-first and agent-native gives me hope. Because those products are much more fun for me to build. And their impact on our civilization may well be much higher, since agent-native usage is starting to overtake human-first usage as we speak.
(If only we could fix banking and taxes and international travel and visas and airport document checks same way, with APIs and agents and strong contracts. What a world that would be. Well, a man can dream.)
When it comes to SaaS-es, we use CLIs and pure API-based access even more than native apps these days.
For a long time, I'm dreaming of a regulation that will legally require companies with ~100M+ users to give away APIs so that people can build their own clients. Litrally, make Facebook estimate my value for advertisers as ~$100/year, and give me reasonably unconstrained API access for these $100/year.
If Meta says I'm worth $10K a year to advertisers, well, let's talk โ that's what Congress hearings and the Attorney General are for. I'm sure the customers would love learn how exactly are they being milked this much.
If Meta says I'm worth $10 a year, well, even better for me.
And then there's an open market for "third-party" clients. Which the very Meta-build Facebook App competes with, freely.
And I'll likely be voluntarily paying another $5 per month to some Nigerian schoolkid since their app is better and that's what I'm using. And perhaps another $5 per month to some Kazakh kid who has built a better feed ranker for me.
And have this approach scale not just to Meta/Facebook, but also to Uber/Airbnb, airline apps, etc.
Well, not going to happen. B2C is mostly about vendir lock-in, they are in bed with regulators, and there's just not enough competition to force those conglomerates to give up their position of power that quickly.
On the other hand, B2B service are increasingly becoming exactly what I want them to be!
Github โ I mostly use it from my agents. YouTrack, because that's what we use โ agent-first usage. Slack and email and calendar โ MCP and here we go.
In fact, for YouTrack I even have my agents use it via the API, not the CLI. The annoying CLI always tries to make me unlock my keychain, making MacOS more like Windows โ annoying af with those pop-ups. But via the API, with the key in an
.env file of that one repo I open with Cursor to triage tasks โ fantastic.Not sure there is a deep moral to this post. My personal life is still about B2C apps, and whatever I can't stand in, say, Google Maps or in Facebook โ that's unlikely to change any time soon.
But that B2B services are increasingly API-first and agent-native gives me hope. Because those products are much more fun for me to build. And their impact on our civilization may well be much higher, since agent-native usage is starting to overtake human-first usage as we speak.
(If only we could fix banking and taxes and international travel and visas and airport document checks same way, with APIs and agents and strong contracts. What a world that would be. Well, a man can dream.)
๐ฅ2
Why is he Jean Reno and not Jean Renault?
๐7
Visualizations are the best way for a human to review a large code change.
I converged on this several weeks ago, but somehow never wrote about it.
Here's the simple mental model. The human brain is ultra-powerful compared to the model. Unless you're doing some very advanced low-level machinery, though, that power shows up mainly when it comes to the big picture.
So, to review a 10K+ LOC pull request, or an entirely new feature, the only way to use the human brain efficiently is to present that big picture to it.
Of course, this will not replace ordinary code reviews. Bugs do creep in from time to time. And various agents โ humans and AIs โ are good at digging deep and hunting them down. But, all in all, those are details.
Simply put, I believe agents are already good enough that, in a good engineering team today, the density of bugs per feature โ or per line of code โ is actually lower than in most software products shipped over the past few decades. So our code is not perfect, but local bugs are certainly not a problem per se.
What is a problem is complexity creep and architecture slop. Agents appear to endorse it, and even, to a certain degree, long for it. They are too bureaucratic, and too focused on showing some result, no matter how ugly things are behind the scenes.
And it's this ugliness that we, humans, have to protect our codebases against. My CEO calls it "architecture slop," and it's a good term.
Figuring out what is essential and what is architecture slop is no easy task. For every precisely targeted question, the agent will have a perfect answer. And reading all the code is already infeasible.
So my solution here is twofold. First, best architectural practices. Second, visualizations.
For best architectural practices, the dataflow-first, events-first, CQRS + CRDT + actor-model way of thinking has never let me down. I need to understand what follows from what. What downstream events can asynchronously affect upstream components. What the acceptable and unacceptable not-quite-right states of the data are. What the failure modes are. And when I ask the model to explain the dataflow from this angle, and to write targeted tests focused on component isolation and logical ordering, things do come together quite nicely.
But visualization is separate, and it's a damn superpower. If a feature takes ~days, then allocating ~hours to have it plot itself as a nice diagram is far too revealing to pass up. Especially if this diagram is code-first โ as in, not a description of what is ๐๐๐๐๐ก to be, but a description of what it ๐๐ . And doubly so if the diagram is โ๐ฆ๐๐๐๐ก๐๐ โ that is, it shows how real data, usually from tests or benchmarks, flows through the system, with interactive drill-down capabilities.
I'm not sure I can explain this well in text, and all the examples I have today are proprietary. I'll share one right away as soon as something open comes up.
The moral is that good engineering and architecture practices are not dead. They remain extremely useful. We just need to apply them inward โ to find better ways for us, humans, to understand and manage the complexity of what we are building and shipping as we speak.
I converged on this several weeks ago, but somehow never wrote about it.
Here's the simple mental model. The human brain is ultra-powerful compared to the model. Unless you're doing some very advanced low-level machinery, though, that power shows up mainly when it comes to the big picture.
So, to review a 10K+ LOC pull request, or an entirely new feature, the only way to use the human brain efficiently is to present that big picture to it.
Of course, this will not replace ordinary code reviews. Bugs do creep in from time to time. And various agents โ humans and AIs โ are good at digging deep and hunting them down. But, all in all, those are details.
Simply put, I believe agents are already good enough that, in a good engineering team today, the density of bugs per feature โ or per line of code โ is actually lower than in most software products shipped over the past few decades. So our code is not perfect, but local bugs are certainly not a problem per se.
What is a problem is complexity creep and architecture slop. Agents appear to endorse it, and even, to a certain degree, long for it. They are too bureaucratic, and too focused on showing some result, no matter how ugly things are behind the scenes.
And it's this ugliness that we, humans, have to protect our codebases against. My CEO calls it "architecture slop," and it's a good term.
Figuring out what is essential and what is architecture slop is no easy task. For every precisely targeted question, the agent will have a perfect answer. And reading all the code is already infeasible.
So my solution here is twofold. First, best architectural practices. Second, visualizations.
For best architectural practices, the dataflow-first, events-first, CQRS + CRDT + actor-model way of thinking has never let me down. I need to understand what follows from what. What downstream events can asynchronously affect upstream components. What the acceptable and unacceptable not-quite-right states of the data are. What the failure modes are. And when I ask the model to explain the dataflow from this angle, and to write targeted tests focused on component isolation and logical ordering, things do come together quite nicely.
But visualization is separate, and it's a damn superpower. If a feature takes ~days, then allocating ~hours to have it plot itself as a nice diagram is far too revealing to pass up. Especially if this diagram is code-first โ as in, not a description of what is ๐๐๐๐๐ก to be, but a description of what it ๐๐ . And doubly so if the diagram is โ๐ฆ๐๐๐๐ก๐๐ โ that is, it shows how real data, usually from tests or benchmarks, flows through the system, with interactive drill-down capabilities.
I'm not sure I can explain this well in text, and all the examples I have today are proprietary. I'll share one right away as soon as something open comes up.
The moral is that good engineering and architecture practices are not dead. They remain extremely useful. We just need to apply them inward โ to find better ways for us, humans, to understand and manage the complexity of what we are building and shipping as we speak.
๐ฅ2โค1
AI is not "ruining" human civilization. AI is just a major catalyst that shows what exactly has already been broken for a long time โ while we keep pretending it's working as intended.
I've been looking for the right moment to write this down, and it feels like the moment has come. What follows is part sentiment, part illustration โ so bear with me while I walk through a few examples before getting to what actually prompted it.
Take tax audits, for instance. Clearly, last year's models are more than capable of finding massive gaps, detecting fraud, and saving our budgets insane amounts of money. However, I very much doubt that we'll see major cuts in fraud in the next few years. Because the incentives are just too misaligned.
Or take patents. The originally advertised idea was and remains to motivate inventors to invent โ since a good invention can, and should, presumably, make one well off. The brutal reality is that patents are used for many purposes, but most definitely not to protect the interests of inventors.
Or monopolies and anti-monopoly committees. They have been defunct for a long time, and we all know full well what legal entities (and what individuals) are profiting tremendously because those monopolies do exist and are far too powerful โ and yet nothing is being done. Those "huge" fines are a rounding error next to how much those entities are making from the very fact that others are effectively deprived of choice.
Ah, and my long-time favorite: science. Peer reviews, scientific consensus, attribution, who gets the credit, who gets the prize. Just ... come on. We all know this facet of the world has been broken for a long, long time. It doesn't take much digging into history to learn how tobacco companies lobbied for the idea that "all opinions must be considered" when it comes to whether smoking causes lung cancer. This has cost lives โ millions of lives. We know who did it, and we know who paid for it. Smoking did decline, eventually, and the settlements were paid. But the playbook itself โ manufacture doubt, demand "balance" โ was never retired. It is still in use today, on other topics.
This is not to mention the years Andrew Wiles spent working in secrecy. No dirty secret there, of course โ but even the most notable mathematical proofs by humanity show that attribution was and remains a cornerstone problem. Incompatible with the real Socratic scientific method, to my taste, but such is life.
And today, suddenly, we are supposed to be outraged about potential misattribution of the Navier-Stokes solution. Solutions, in the plural, I should say โ if any of them holds. My bet is that quite a few claims from it will hold, although practical applications of that proof remain to be seen.
I can make a broader point here, but it will be too bleak. So let me leave it as is: just don't forget that AI is not "destroying" the "fabric of our reality". AI is just showing where we chose not to look up for far too long.
And, despite being 40+, I still consider myself young enough to believe this is ultimately a good thing. Because the more we do look up, the higher the chance of us, humankind, not doing something very stupid collectively.
I've been looking for the right moment to write this down, and it feels like the moment has come. What follows is part sentiment, part illustration โ so bear with me while I walk through a few examples before getting to what actually prompted it.
Take tax audits, for instance. Clearly, last year's models are more than capable of finding massive gaps, detecting fraud, and saving our budgets insane amounts of money. However, I very much doubt that we'll see major cuts in fraud in the next few years. Because the incentives are just too misaligned.
Or take patents. The originally advertised idea was and remains to motivate inventors to invent โ since a good invention can, and should, presumably, make one well off. The brutal reality is that patents are used for many purposes, but most definitely not to protect the interests of inventors.
Or monopolies and anti-monopoly committees. They have been defunct for a long time, and we all know full well what legal entities (and what individuals) are profiting tremendously because those monopolies do exist and are far too powerful โ and yet nothing is being done. Those "huge" fines are a rounding error next to how much those entities are making from the very fact that others are effectively deprived of choice.
Ah, and my long-time favorite: science. Peer reviews, scientific consensus, attribution, who gets the credit, who gets the prize. Just ... come on. We all know this facet of the world has been broken for a long, long time. It doesn't take much digging into history to learn how tobacco companies lobbied for the idea that "all opinions must be considered" when it comes to whether smoking causes lung cancer. This has cost lives โ millions of lives. We know who did it, and we know who paid for it. Smoking did decline, eventually, and the settlements were paid. But the playbook itself โ manufacture doubt, demand "balance" โ was never retired. It is still in use today, on other topics.
This is not to mention the years Andrew Wiles spent working in secrecy. No dirty secret there, of course โ but even the most notable mathematical proofs by humanity show that attribution was and remains a cornerstone problem. Incompatible with the real Socratic scientific method, to my taste, but such is life.
And today, suddenly, we are supposed to be outraged about potential misattribution of the Navier-Stokes solution. Solutions, in the plural, I should say โ if any of them holds. My bet is that quite a few claims from it will hold, although practical applications of that proof remain to be seen.
I can make a broader point here, but it will be too bleak. So let me leave it as is: just don't forget that AI is not "destroying" the "fabric of our reality". AI is just showing where we chose not to look up for far too long.
And, despite being 40+, I still consider myself young enough to believe this is ultimately a good thing. Because the more we do look up, the higher the chance of us, humankind, not doing something very stupid collectively.
๐10โค5
To add to the previous post: https://x.com/ValerioCapraro/status/2097791836269977996
We were supposed to act surprised when Google settled a class action lawsuit related to deleting ("not retaining"?) Incognito window logs it kept for a long time. So much for Don't be Evil.
And now we are supposed to believe OpenAI et. al. would not use whatever data they can get a hold of during training and/or reasoning.
Seriously, this may well be a fluke originally. Something like nginx logs not properly rotated. And maybe a good-faith kind-hearted pro-privacy SRE said "wait, these are good for debugging, but it appears our models are using this, so let's perhaps restrict access, or out right delete these logs?"
And then someone From The Business stepped in and said, or rather, quietly without-saying-anything, made sure those logs do stay and do remain available for the model.
And this other someone From The Business probably got paid ~1000x more than the diligent SRE who sincerely wanted to set the record straight.
I am not saying this is what happened. It's totally made up by yours truly. Any and every relation to reality is pure coincidence.
But show me the incentives and I'll show you the outcome. That's how the world works. And if you think otherwise, it's probably about time to grow up.
We were supposed to act surprised when Google settled a class action lawsuit related to deleting ("not retaining"?) Incognito window logs it kept for a long time. So much for Don't be Evil.
And now we are supposed to believe OpenAI et. al. would not use whatever data they can get a hold of during training and/or reasoning.
Seriously, this may well be a fluke originally. Something like nginx logs not properly rotated. And maybe a good-faith kind-hearted pro-privacy SRE said "wait, these are good for debugging, but it appears our models are using this, so let's perhaps restrict access, or out right delete these logs?"
And then someone From The Business stepped in and said, or rather, quietly without-saying-anything, made sure those logs do stay and do remain available for the model.
And this other someone From The Business probably got paid ~1000x more than the diligent SRE who sincerely wanted to set the record straight.
I am not saying this is what happened. It's totally made up by yours truly. Any and every relation to reality is pure coincidence.
But show me the incentives and I'll show you the outcome. That's how the world works. And if you think otherwise, it's probably about time to grow up.
X (formerly Twitter)
Valerio Capraro (@ValerioCapraro) on X
BREAKING: OpenAI might have stolen another major proof.
In a detailed Mastodon post, which I report in full in the comments, Andreas Thom presents several pieces of evidence suggesting that Openโฆ
In a detailed Mastodon post, which I report in full in the comments, Andreas Thom presents several pieces of evidence suggesting that Openโฆ
๐8
This post was deleted by WEF.
"You can't make this up" (c)
Claude:
"You can't make this up" (c)
Claude:
Yes. As for deletions, WEF has been scrubbing this stuff in stages for years. In November 2020, WEF deleted the "8 predictions" video after "You'll own nothing and be happy" started trending on Twitter. By 2022, the essay was no longer available on WEF's website, and a June 2022 blog post claimed WEF deleted the "you will own nothing" tweet. The "8 Predictions" article itself was taken down from WEF's site around AprilโMay 2023.
๐2
Thoughts after sleeping on a recent incident where a major bank handed its customers' passports, selfies, and transaction histories to scammers who asked from a government email address. I can now phrase my position even stronger.
TL;DR: The problem is not just what the regulators and the governments are doing, although they are doing a lot. The problem is that they are too powerful, and no one can stand up to them and say "enough!"
This was not the case ten years ago. It was, five years ago already. I've seen it in DeFi / Web3, and it's just too obvious with ordinary banking.
Banks want to do business. It's too good of a business to be a bank. And to be a bank you have to be friends with the regulator. Regulators, in turn, want more and more bank friends, as long as they play by the rules. And the banks do. Of course they do.
What we lack is a force on the side of We, The People strong enough that some bank cracks under the pressure and openly states:
โข Screw you, governments and regulators.
โข Your rules are NOT protecting the customer, they are making it WORSE.
โข You want more control, but it's just PRIVACY INVASIONS, and they WILL backfire.
โข So yes, we, bank X, ARE WILLING to serve customers in your jurisdiction.
โข But we will open OFFSHORE accounts for them, and ship cards by international mail. Perhaps under different names too, which is perfectly legal: the name on your card need not match your passport.
โข Not to evade your rules. To PROTECT our customers.
โข And now, dear regulator, PLEASE WORK WITH US. We can help you draft good regulation and establish guardrails, so that no regulator, YOU TODAY OR THE NEXT ONE TOMORROW, could stop putting bank clients' interests first.
There is literally NO SUCH BANK. And this is the problem.
I'm sure tons of founders in banking would LOVE to say this. They are already billionaires, so if such a move had a real chance, I would expect someone to have tried it by now. But the governments and regulators are too damn powerful. Power corrupts, and absolute power corrupts absolutely. When it comes to financial overseeing, our overlords are corrupted.
And I am saying this as an ordinary, law-abiding citizen. I do support KYC and AML. I pass them regularly. I want more people to pass them when buying a house next to mine, a yacht, or a large stake in a startup. Hell yeah, track the source of income.
But the system should be designed bottom up. Not "we, the Regulator, must see everything", but "here is the list of acceptable documentation: provide it and you are GUARANTEED to pass KYC/AML". Tax forms and income statements, for instance. Not one's home address, for God's sake, not one's net worth, and definitely not one's family and family businesses.
A trivial change, really. And finally: make it unconditionally legal to bank in a different jurisdiction. Today, banks in plenty of countries refuse citizens of certain other countries, only because those countries make it too damn hard. It should be perfectly fine, as long as one pays their taxes at home. Fierce competition would follow, and transfers would take seconds for EVERYBODY, because commerce works. Top-down suffocation does not.
The Web3 folks have already built the rails. Seconds-long settlements for fractions of a penny, at scale, with receipts and proofs attachable to any transaction, so that the parties CAN CHOOSE to show them, decrypt them, or ZK-prove something about them without revealing the rest. With an on-chain USD[C] account and Web3-friendly payroll and brokers, one's tax forms are a) fully automatic, b) perfectly auditable, and c) impossible to make a mistake in. Yes, impossible. Some operations, designed right, leave no room for a mistake, and getting paid while the tax authority knows about it is one of them.
We had a chance to live in this world by 2024. Or so it seemed in 2016. We blew it up.
But if enough sane powerful people say "enough!", the tide may well turn. That is what I would love to see, as an ordinary user of the world's banking system.
TL;DR: The problem is not just what the regulators and the governments are doing, although they are doing a lot. The problem is that they are too powerful, and no one can stand up to them and say "enough!"
This was not the case ten years ago. It was, five years ago already. I've seen it in DeFi / Web3, and it's just too obvious with ordinary banking.
Banks want to do business. It's too good of a business to be a bank. And to be a bank you have to be friends with the regulator. Regulators, in turn, want more and more bank friends, as long as they play by the rules. And the banks do. Of course they do.
What we lack is a force on the side of We, The People strong enough that some bank cracks under the pressure and openly states:
โข Screw you, governments and regulators.
โข Your rules are NOT protecting the customer, they are making it WORSE.
โข You want more control, but it's just PRIVACY INVASIONS, and they WILL backfire.
โข So yes, we, bank X, ARE WILLING to serve customers in your jurisdiction.
โข But we will open OFFSHORE accounts for them, and ship cards by international mail. Perhaps under different names too, which is perfectly legal: the name on your card need not match your passport.
โข Not to evade your rules. To PROTECT our customers.
โข And now, dear regulator, PLEASE WORK WITH US. We can help you draft good regulation and establish guardrails, so that no regulator, YOU TODAY OR THE NEXT ONE TOMORROW, could stop putting bank clients' interests first.
There is literally NO SUCH BANK. And this is the problem.
I'm sure tons of founders in banking would LOVE to say this. They are already billionaires, so if such a move had a real chance, I would expect someone to have tried it by now. But the governments and regulators are too damn powerful. Power corrupts, and absolute power corrupts absolutely. When it comes to financial overseeing, our overlords are corrupted.
And I am saying this as an ordinary, law-abiding citizen. I do support KYC and AML. I pass them regularly. I want more people to pass them when buying a house next to mine, a yacht, or a large stake in a startup. Hell yeah, track the source of income.
But the system should be designed bottom up. Not "we, the Regulator, must see everything", but "here is the list of acceptable documentation: provide it and you are GUARANTEED to pass KYC/AML". Tax forms and income statements, for instance. Not one's home address, for God's sake, not one's net worth, and definitely not one's family and family businesses.
A trivial change, really. And finally: make it unconditionally legal to bank in a different jurisdiction. Today, banks in plenty of countries refuse citizens of certain other countries, only because those countries make it too damn hard. It should be perfectly fine, as long as one pays their taxes at home. Fierce competition would follow, and transfers would take seconds for EVERYBODY, because commerce works. Top-down suffocation does not.
The Web3 folks have already built the rails. Seconds-long settlements for fractions of a penny, at scale, with receipts and proofs attachable to any transaction, so that the parties CAN CHOOSE to show them, decrypt them, or ZK-prove something about them without revealing the rest. With an on-chain USD[C] account and Web3-friendly payroll and brokers, one's tax forms are a) fully automatic, b) perfectly auditable, and c) impossible to make a mistake in. Yes, impossible. Some operations, designed right, leave no room for a mistake, and getting paid while the tax authority knows about it is one of them.
We had a chance to live in this world by 2024. Or so it seemed in 2016. We blew it up.
But if enough sane powerful people say "enough!", the tide may well turn. That is what I would love to see, as an ordinary user of the world's banking system.
โค5๐1๐ข1
I just love working with modern-day agents, and I found a new pattern I want to share.
Here it is in one sentence: I tell the agent on my laptop to configure a remote machine for me, and then to go do the actual work THERE, inside tmux, driving a second agent on that remote box. The laptop agent sets it up, the laptop agent uses it, and I never touch the remote machine by hand โ besides creating a mundane, unprivileged user and configuring their SSH access. That, plus the occasional sudo command later on, is the only elevated work that box ever sees, and I do it myself. Everything else happens as that user, and it happens through agents.
Now, the rule above all: my sensitive credentials never leave the laptop, and every single operation that needs them requires Touch ID. Root anywhere, my main GitHub account, anything that can do real damage: Touch ID, every time. The agents' SSH key to the remote box is NOT one of those. That key is free; agents can and should use it as they please, and the worst it gets them is a shell as a user with no root. So an agent cannot suddenly open a root session, locally or on that machine. I would have to be here to approve it, and I sure as hell will not approve a surprise. When something on the remote machine does require sudo, I ask my agents to tell me what the commands are, and I run them one by one, manually, after I authenticate with Touch ID.
Then I add a Host entry for the machine to ~/.ssh/config on the laptop, under a random name that I also save under /etc/hosts. So that
And that's where the human part ends. I already have a skill on my laptop that configures a remote machine and its user, without root, into whatever my agents need: it installs the tools, sets up tmux, and copies the logins that box needs, which means the fenced GitHub account from the next paragraph and my coding-agent subscriptions. What the experiment itself needs are the .env files from the repository, and those are perfectly fine to share; that is the entire point of an isolated remote user. My sensitive credentials never leave the laptop. I use Podman instead of Docker to run tasks there. In a few minutes, from a fresh box, it's a beast.
On those coding-agent subscriptions: yes, I log my own subscriptions into that box. It is still just me using them, through my own agent from another machine, and I've triple-checked that this is within their terms. Having to re-log some of them in every once in a while is a tad annoying, but, quite frankly, it's a low price to pay for the productivity boost and for the overall experience.
GitHub is fenced the same way. The remote machine gets a separate GitHub account that cannot push into anything I consider important โ that would require an approval from my laptop, and my Touch ID, via the Passkey / WebAuthn protocol. It can read the repositories I've granted it access to. It can comment, make commits, and open pull requests. That's it. Merging stays with me.
Here it is in one sentence: I tell the agent on my laptop to configure a remote machine for me, and then to go do the actual work THERE, inside tmux, driving a second agent on that remote box. The laptop agent sets it up, the laptop agent uses it, and I never touch the remote machine by hand โ besides creating a mundane, unprivileged user and configuring their SSH access. That, plus the occasional sudo command later on, is the only elevated work that box ever sees, and I do it myself. Everything else happens as that user, and it happens through agents.
Now, the rule above all: my sensitive credentials never leave the laptop, and every single operation that needs them requires Touch ID. Root anywhere, my main GitHub account, anything that can do real damage: Touch ID, every time. The agents' SSH key to the remote box is NOT one of those. That key is free; agents can and should use it as they please, and the worst it gets them is a shell as a user with no root. So an agent cannot suddenly open a root session, locally or on that machine. I would have to be here to approve it, and I sure as hell will not approve a surprise. When something on the remote machine does require sudo, I ask my agents to tell me what the commands are, and I run them one by one, manually, after I authenticate with Touch ID.
Then I add a Host entry for the machine to ~/.ssh/config on the laptop, under a random name that I also save under /etc/hosts. So that
ssh something just takes me โ or my agent โ straight there. The same entry forwards a few ports, so while a session is open I can hit whatever is running on the remote box straight from my laptop. I tell the laptop agent: this machine exists, this is its name, and these ports are yours as long as the connection remains open. The agent manages that connection while we are working together; not me. When the connection closes, the ports go away, but the work does not: it lives in tmux on the remote box, and the next connection picks it right back up. This, by the way, is something I'm thinking of improving over time, but so far it is good enough and it has not let me down.And that's where the human part ends. I already have a skill on my laptop that configures a remote machine and its user, without root, into whatever my agents need: it installs the tools, sets up tmux, and copies the logins that box needs, which means the fenced GitHub account from the next paragraph and my coding-agent subscriptions. What the experiment itself needs are the .env files from the repository, and those are perfectly fine to share; that is the entire point of an isolated remote user. My sensitive credentials never leave the laptop. I use Podman instead of Docker to run tasks there. In a few minutes, from a fresh box, it's a beast.
On those coding-agent subscriptions: yes, I log my own subscriptions into that box. It is still just me using them, through my own agent from another machine, and I've triple-checked that this is within their terms. Having to re-log some of them in every once in a while is a tad annoying, but, quite frankly, it's a low price to pay for the productivity boost and for the overall experience.
GitHub is fenced the same way. The remote machine gets a separate GitHub account that cannot push into anything I consider important โ that would require an approval from my laptop, and my Touch ID, via the Passkey / WebAuthn protocol. It can read the repositories I've granted it access to. It can comment, make commits, and open pull requests. That's it. Merging stays with me.
๐3
Now the fun part. I take a piece of work that is semi-complete, or not isolated properly, or just too heavy for my laptop. (Yes, too heavy does happen often, I have drained my MacBook Pro's battery 100% to 3% within just ~1.5 hours several times, and I do not like how warm it gets.) Then I tell my laptop agent โ let's say it's Cursor because it often is โ "This should be done on that machine, in a different harness. Go there, open a tmux session, copy whatever you need to make it happen. Create the environment from the one in this repository; it's .gitignore-d because it is machine-local, not because it is secret, so you are allowed to share it. Run it there, and tell me how to check on it later."
And it does. The laptop agent SSHes in, sets up the session, launches the agent on the remote box inside tmux, and disconnects. tmux is what makes this work: the remote session outlives the connection, and it turns out remote agents can drive tmux perfectly well via send-keys, so I am just a happy observer here. There is no live laptop session to keep alive and no babysitting. I can close the laptop, reboot it, come back the next day, and ask the agent to check on the work and clean up. It does not drain my battery. It is practically free for the laptop.
What I get out of it: real isolation from my laptop and from my sensitive credentials, because the remote user has none of them. And a setup that is cheap to recreate. Once the work is done, I ask the agent on either machine to document everything and commit it. Then I create a NEW user on that box and repeat the whole thing from scratch, to prove it actually reproduces.
Also, a shameless plug: my Scoped Skills Helper,
Ain't this cool?
And it does. The laptop agent SSHes in, sets up the session, launches the agent on the remote box inside tmux, and disconnects. tmux is what makes this work: the remote session outlives the connection, and it turns out remote agents can drive tmux perfectly well via send-keys, so I am just a happy observer here. There is no live laptop session to keep alive and no babysitting. I can close the laptop, reboot it, come back the next day, and ask the agent to check on the work and clean up. It does not drain my battery. It is practically free for the laptop.
What I get out of it: real isolation from my laptop and from my sensitive credentials, because the remote user has none of them. And a setup that is cheap to recreate. Once the work is done, I ask the agent on either machine to document everything and commit it. Then I create a NEW user on that box and repeat the whole thing from scratch, to prove it actually reproduces.
Also, a shameless plug: my Scoped Skills Helper,
scsh, is what I use on that remote machine to orchestrate swarms of agents that are run under different harnesses. Different parts of the work are done by different model providers, and I have a complete asciinema recording of every single process, with annotated timestamps of what happened when, and how many tokens each invocation took.Ain't this cool?
๐ฅ6