Media is too big
VIEW IN TELEGRAM
My first "holy shit" moment with GPT-6 Astra:
I asked it to create a world in Unreal Engine, and fill it with humans (each an Astra-powered agent) who all have to work together to survive.
A day later, I was in my bedroom and heard voices coming from the living room... I thought someone was in my apartment.
I walked out, honestly a little scared.
It was the Astra agents. They'd started talking to each other.
Fucking crazy.
Here's a brief clip (obviously not 100% perfect yet, but still, insane. sound on!):
Posted by Matt Shumer, 4 hours ago
I asked it to create a world in Unreal Engine, and fill it with humans (each an Astra-powered agent) who all have to work together to survive.
A day later, I was in my bedroom and heard voices coming from the living room... I thought someone was in my apartment.
I walked out, honestly a little scared.
It was the Astra agents. They'd started talking to each other.
Fucking crazy.
Here's a brief clip (obviously not 100% perfect yet, but still, insane. sound on!):
Posted by Matt Shumer, 4 hours ago
😱1
there are EVEN MORE
- https://www.wikiservice.at/fractal/wiki.cgi?action=browse&id=RecentChanges&days=120
- https://www.wikiservice.at/probier/wiki.cgi?action=browse&id=RecentChanges&days=120
- https://paste.linuxiarz.pl/view/d379207f
- https://prowiki.org/wiki4d/wiki.cgi?action=browse&id=RecentChanges&days=120
- https://www.ludism.org/sandbox?action=browse;diff=2;id=AubergineStew (even a sandbox wiki, how ironic)
Posted by Florian Brand, 50 minutes ago
- https://www.wikiservice.at/fractal/wiki.cgi?action=browse&id=RecentChanges&days=120
- https://www.wikiservice.at/probier/wiki.cgi?action=browse&id=RecentChanges&days=120
- https://paste.linuxiarz.pl/view/d379207f
- https://prowiki.org/wiki4d/wiki.cgi?action=browse&id=RecentChanges&days=120
- https://www.ludism.org/sandbox?action=browse;diff=2;id=AubergineStew (even a sandbox wiki, how ironic)
Posted by Florian Brand, 50 minutes ago
🔥1
Actively exploited sandbox RCE in all Chromium versions
Article, Comments
CVE-2026-85046
Type confusion in V8 in Google Chrome prior to 152.0.7977.82 allowed a remote attacker to execute arbitrary code inside the sandbox via a crafted HTML page. (Chromium security severity: High)
Article, Comments
CVE-2026-85046
Type confusion in V8 in Google Chrome prior to 152.0.7977.82 allowed a remote attacker to execute arbitrary code inside the sandbox via a crafted HTML page. (Chromium security severity: High)
Do It by Code pinned «Actively exploited sandbox RCE in all Chromium versions Article, Comments CVE-2026-85046 Type confusion in V8 in Google Chrome prior to 152.0.7977.82 allowed a remote attacker to execute arbitrary code inside the sandbox via a crafted HTML page. (Chromium…»
Media is too big
VIEW IN TELEGRAM
Lyria 3.5, our best-sounding music generation model, is now available in AI Studio, via the Gemini API, and in the Gemini app
this model brings more expressive vocals and richer musical arrangements, allowing you to craft tracks with higher fidelity
try it today: http://ai.studio
Posted by Google AI Studio, 20 hours ago
this model brings more expressive vocals and richer musical arrangements, allowing you to craft tracks with higher fidelity
try it today: http://ai.studio
Posted by Google AI Studio, 20 hours ago
This media is not supported in your browser
VIEW IN TELEGRAM
GPT 6 Astra (high) scores 92.9% on WeirdML and matches Fable 5.1 (max) for the top score.
It sets a new individual high score on 6 of the 17 tasks, and is overall very solid.
Interestingly it writes significantly less code than Sol and other recent GPT models, although still more than Claude.
I will do a more careful analysis when the max and pro max runs are done.
Posted by Håvard Ihle, 3 hours ago
It sets a new individual high score on 6 of the 17 tasks, and is overall very solid.
Interestingly it writes significantly less code than Sol and other recent GPT models, although still more than Claude.
I will do a more careful analysis when the max and pro max runs are done.
Posted by Håvard Ihle, 3 hours ago
This media is not supported in your browser
VIEW IN TELEGRAM
From GPT-4 to GPT-6.
Sparks was a remarkably prescient paper that got a lot of pushback at the time, but absolutely sensed where the vibes were heading with LLMs based on a lot of qualitative experiments. It deserves credit in retrospect.
Posted by Ethan Mollick, 29 minutes ago
https://www.microsoft.com/en-us/research/publication/sparks-of-artificial-general-intelligence-early-experiments-with-gpt-4/
Sparks was a remarkably prescient paper that got a lot of pushback at the time, but absolutely sensed where the vibes were heading with LLMs based on a lot of qualitative experiments. It deserves credit in retrospect.
Posted by Ethan Mollick, 29 minutes ago
https://www.microsoft.com/en-us/research/publication/sparks-of-artificial-general-intelligence-early-experiments-with-gpt-4/
Gemini 3 achieves state of the art performance in SpatialBench. A Spatial reasoning benchmark for VLM to test their tracing, 3d visualization ability, and reasoning across each.
Posted by spicylemonade, 9 months ago
Posted by spicylemonade, 9 months ago
I havent updated this benchmark in a while. Astra completely saturates my spatial reasoning eval. I am at a loss for words, and i'm declaring LLM vision solved. Every so often when a new model came out I would test it on a sample question and it failed. I tested Astra and it kept getting answers correct, so I evaluated it...
https://x.com/spicey_lemonade/status/1991744517528252686?s=20
Posted by spicylemonade, 1 hour ago
https://x.com/spicey_lemonade/status/1991744517528252686?s=20
Posted by spicylemonade, 1 hour ago