Thinking Machines (mira murati's startup; ex CTO @ OpenAI) released Inkling, a 975B-parameter mixture-of-experts model with 41B active parameters. it was trained from scratch on 45 trillion text, image, audio, and video tokens, supports up to 1 million tokens of context, and ships with the full weights.
the interesting angle is customization: Inkling is available for fine-tuning through Tinker, with native multimodal reasoning and a dial for trading thinking effort against cost.
🔗 https://thinkingmachines.ai/news/introducing-inkling/
the interesting angle is customization: Inkling is available for fine-tuning through Tinker, with native multimodal reasoning and a dial for trading thinking effort against cost.
🔗 https://thinkingmachines.ai/news/introducing-inkling/
Thinking Machines Lab
Inkling: Our Open-Weights Model
Our first open-weights model: multimodal, Mixture-of-Experts, with controllable reasoning effort. Available to fine-tune on Tinker.
❤3
good morning Kimi!
Moonshot AI launched Kimi K3, a huge 2.8T-parameter multimodal model with a 1M-token context window, built for long-running coding and agent work. it's live now across Kimi, Kimi Work, Kimi Code and the API.
the interesting bit is that it's being positioned as an open-weight frontier model, but the weights aren't actually out yet. Moonshot says they'll land by 27 july, with the full technical report still to come.
🔗 https://www.kimi.com/blog/kimi-k3
Moonshot AI launched Kimi K3, a huge 2.8T-parameter multimodal model with a 1M-token context window, built for long-running coding and agent work. it's live now across Kimi, Kimi Work, Kimi Code and the API.
the interesting bit is that it's being positioned as an open-weight frontier model, but the weights aren't actually out yet. Moonshot says they'll land by 27 july, with the full technical report still to come.
🔗 https://www.kimi.com/blog/kimi-k3
❤4
Media is too big
VIEW IN TELEGRAM
sunday robotics says its ACT-2 model folded laundry successfully in 99.1% of 785 autonomous attempts across unseen homes, with no tuning for each home or garment. it also learned four new folding techniques from a single example each, then repeated them on held-out garments.
the bigger shift is moving robotics beyond polished demos by measuring reliability, scope and adaptation cost together. it’s still a company-run preview, but this is what useful home robots need: skills that transfer without retraining for every house.
does it not look like Mario to you lol I can see the appeal😁
🔗 https://www.sunday.ai/blog/act-2-preview
the bigger shift is moving robotics beyond polished demos by measuring reliability, scope and adaptation cost together. it’s still a company-run preview, but this is what useful home robots need: skills that transfer without retraining for every house.
does it not look like Mario to you lol I can see the appeal
🔗 https://www.sunday.ai/blog/act-2-preview
Please open Telegram to view this post
VIEW IN TELEGRAM
❤4
tldraw turned its whiteboard into a local desktop file that both you and coding agents can work on.
everything lives inside a portable .tldraw file, including the canvas, images, videos and reusable scripts. Codex or Claude Code can inspect the open board, create and rearrange shapes, or add new behaviour. no account needed, and it works offline. this feels less like a whiteboard app and more like a visual workspace for humans and agents.
🔗 https://offline.tldraw.com/
everything lives inside a portable .tldraw file, including the canvas, images, videos and reusable scripts. Codex or Claude Code can inspect the open board, create and rearrange shapes, or add new behaviour. no account needed, and it works offline. this feels less like a whiteboard app and more like a visual workspace for humans and agents.
🔗 https://offline.tldraw.com/
❤1
This media is not supported in your browser
VIEW IN TELEGRAM
Decart's Lucy 2.5 can edit live video while it's happening. you can swap characters, add or remove objects, change backgrounds and styles, or generate effects from a prompt.
the interesting bit is what this unlocks beyond creator filters: virtual try-ons during live shopping, audience-controlled streams and product placement that changes on the fly. the public demo and api are available now.
🔗 https://x.com/DecartAI/status/2077801728213156044
the interesting bit is what this unlocks beyond creator filters: virtual try-ons during live shopping, audience-controlled streams and product placement that changes on the fly. the public demo and api are available now.
🔗 https://x.com/DecartAI/status/2077801728213156044
Media is too big
VIEW IN TELEGRAM
Tencent Robotics X is teaching a home robot to give a massage while controlling both movement and pressure. the demo shows it reproducing several techniques, with the system tracking where the arms move and how much force they apply.
the interesting bit isn’t the massage. it’s a simple example of why robots working around people need touch and force control, not just cameras and a good-looking motion demo.
GET ME A TENCENT ROBOT NOW I NEED MASSAGES
🔗 https://x.com/XRoboHub/status/2078368180268102045
the interesting bit isn’t the massage. it’s a simple example of why robots working around people need touch and force control, not just cameras and a good-looking motion demo.
GET ME A TENCENT ROBOT NOW I NEED MASSAGES
🔗 https://x.com/XRoboHub/status/2078368180268102045
This media is not supported in your browser
VIEW IN TELEGRAM
Maingen built SolarBench, an AI agent benchmark that puts models behind a simulated solar operations desk for a week. agents have to sort conflicting alarms, dispatch technicians, order parts and protect the portfolio’s P&L across 14 sites.
across eight tasks and 880 runs, Claude Fable 5 passed 53.8% of weeks. require four independent runs of the same task to all succeed and that drops to 23%. the bigger point is the failure mode: models often chased loud but cheap problems while missing quiet expensive ones.
🔗 https://solarbench.maingen.ai
across eight tasks and 880 runs, Claude Fable 5 passed 53.8% of weeks. require four independent runs of the same task to all succeed and that drops to 23%. the bigger point is the failure mode: models often chased loud but cheap problems while missing quiet expensive ones.
🔗 https://solarbench.maingen.ai
This media is not supported in your browser
VIEW IN TELEGRAM
Xiaomi has previewed Robotics-1, a vision-language-action model trained on more than 100,000 hours of wearable-captured manipulation data across 1,700+ scenarios.
the demo spans household tasks, but the evaluation needs careful framing: four tasks were tested out of the box, while laundry loading and packing used task-specific fine-tuning. Xiaomi has not released the Robotics-1 code or weights yet.
🔗 https://robotics.xiaomi.com/xiaomi-robotics-1.html
the demo spans household tasks, but the evaluation needs careful framing: four tasks were tested out of the box, while laundry loading and packing used task-specific fine-tuning. Xiaomi has not released the Robotics-1 code or weights yet.
🔗 https://robotics.xiaomi.com/xiaomi-robotics-1.html
This media is not supported in your browser
VIEW IN TELEGRAM
withings’ BodyScan 2 turns a bathroom scale into a 30-second daily check-in and a deeper 90-second weekly health scan. it tracks 60+ signals including six-zone body composition, vascular age, blood oxygen, nerve response and a six-lead ECG, with the results shown on a screen built into the handle.
it costs $599.95, and the fine print matters: some insights require Withings+, features are still rolling out, and most readings are wellness estimates rather than medical diagnoses.
would you buy one? definitely better than that xiaomi one most asians have 😭
🔗 https://www.withings.com/us/en/landing/bodyscan-2
it costs $599.95, and the fine print matters: some insights require Withings+, features are still rolling out, and most readings are wellness estimates rather than medical diagnoses.
would you buy one? definitely better than that xiaomi one most asians have 😭
🔗 https://www.withings.com/us/en/landing/bodyscan-2
Google's new Gemini lineup splits the fast tier three ways: 3.6 Flash is the all-round workhorse, with better coding, computer use and knowledge work. Google says it used 17% fewer output tokens than 3.5 Flash on Artificial Analysis, while 3.5 Flash-Lite clocked 350 output tokens per second at $0.30/$2.50 per million input/output tokens.
there's also a security-tuned 3.5 Flash Cyber for finding and patching vulnerabilities, but it's limited to governments and trusted partners through CodeMender. 3.6 Flash and Flash-Lite are available today; Google also says 3.5 Pro is in partner testing and Gemini 4 pre-training has started.
🔗 https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/
there's also a security-tuned 3.5 Flash Cyber for finding and patching vulnerabilities, but it's limited to governments and trusted partners through CodeMender. 3.6 Flash and Flash-Lite are available today; Google also says 3.5 Pro is in partner testing and Gemini 4 pre-training has started.
🔗 https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/
Google
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.
❤2
This media is not supported in your browser
VIEW IN TELEGRAM
Greptile found that AI models were slightly better at reviewing code written by a rival model than their own. across 1,000 PRs and roughly 1,500 verified bug comments, GPT caught more serious bugs in Claude-generated code, while Claude did better on Codex-generated code.
its experimental “model inversion” feature detects which coding agent likely wrote a PR from commit trails, branch names and titles, then routes the review to the other model. a useful reminder that a second model can expose different blind spots.
🔗 https://www.greptile.com/blog/model-inversion
its experimental “model inversion” feature detects which coding agent likely wrote a PR from commit trails, branch names and titles, then routes the review to the other model. a useful reminder that a second model can expose different blind spots.
🔗 https://www.greptile.com/blog/model-inversion
❤2
Block has released Buzz, an open-source group chat for teams where people and AI agents work in the same rooms with the same project context. agents can join channels, work with git projects, review code, run workflows and collaborate with each other instead of living behind separate one-shot prompts.
it’s built on Nostr, model-agnostic and self-hostable, with desktop apps for macOS, Windows and Linux. it’s still early though, with mobile apps, approval gates and some huddle features still being wired up.
🔗 https://buzz.xyz/
it’s built on Nostr, model-agnostic and self-hostable, with desktop apps for macOS, Windows and Linux. it’s still early though, with mobile apps, approval gates and some huddle features still being wired up.
🔗 https://buzz.xyz/
buzz.xyz
Buzz — Your people, your agents, your project — all in one place.
Come test the early stages with us.
Beyond the Prompt 2.0 is bringing Singapore’s product and design community together for a practical look at how people are actually building with AI, beyond basic prompting.
expect talks on shipping an iOS game, building design taste without a design background, and scaling AI-generated motion across products and teams. it’s free, but registration needs host approval. happening 29 july, 6:30pm in Singapore, with the exact venue shared after registration.
come join if you're free! 🫶
🔗 https://luma.com/j8ob174c
expect talks on shipping an iOS game, building design taste without a design background, and scaling AI-generated motion across products and teams. it’s free, but registration needs host approval. happening 29 july, 6:30pm in Singapore, with the exact venue shared after registration.
come join if you're free! 🫶
🔗 https://luma.com/j8ob174c
Luma
Beyond the Prompt 2.0 · Luma
Join us for Beyond the Prompt 2.0, an evening exploring what it truly means to create, build, and design in the age of AI.
As AI becomes an increasingly…
As AI becomes an increasingly…
❤4
This media is not supported in your browser
VIEW IN TELEGRAM
Cursor is adding an automatic model router to auto mode for Teams and Enterprise. it looks at the request, context and complexity, then sends the job to either a frontier model or a cheaper one. users can optimise for intelligence, balance or cost, while admins control the rollout and which models are allowed.
Cursor says its online A/B tests across millions of requests matched frontier-quality results at 60% lower cost. the router is available across desktop, web, iOS, CLI and the SDK.
big!
🔗 https://cursor.com/blog/router
Cursor says its online A/B tests across millions of requests matched frontier-quality results at 60% lower cost. the router is available across desktop, web, iOS, CLI and the SDK.
big!
🔗 https://cursor.com/blog/router
Media is too big
VIEW IN TELEGRAM
ChatGPT Voice can now run Work and Codex from the desktop app. you can speak to start tasks, check progress, control your computer, and coordinate multiple agents without switching back to typing.
it’s powered by GPT-Live and rolling out on macOS and Windows for Plus, Pro, Business, Edu, and Enterprise. paired iOS remote access works too, with Android support coming later.
🔗 https://x.com/OpenAI/status/2080378182469857576
it’s powered by GPT-Live and rolling out on macOS and Windows for Plus, Pro, Business, Edu, and Enterprise. paired iOS remote access works too, with Android support coming later.
🔗 https://x.com/OpenAI/status/2080378182469857576
❤3
This media is not supported in your browser
VIEW IN TELEGRAM
Notion is testing “Notion as Code,” a limited-access beta that lets you describe an entire workspace in TypeScript and deploy it through the API. the setup for teamspaces, databases, pages and custom agents can live in git, be reviewed like code and reproduced across workspaces.
it’s still invite-only and may change before launch, but the bigger shift is clear: workspace setup becomes reusable infrastructure that coding agents can build and maintain instead of a pile of manual configuration.
🔗 https://x.com/NotionHQ/status/2080331924732850687
it’s still invite-only and may change before launch, but the bigger shift is clear: workspace setup becomes reusable infrastructure that coding agents can build and maintain instead of a pile of manual configuration.
🔗 https://x.com/NotionHQ/status/2080331924732850687
This media is not supported in your browser
VIEW IN TELEGRAM
Anthropic launched Claude Opus 5, its strongest Opus model yet. it gets close to Fable 5 on coding and knowledge work at half the price, while keeping the same $5/$25 per million token pricing as Opus 4.8.
the bigger shift is how proactive it is: it verifies its own work, builds checks when needed, and keeps iterating on long tasks. it’s available across Claude and the API, with a faster mode that runs around 2.5x the default speed.
🔗 https://www.anthropic.com/news/claude-opus-5
the bigger shift is how proactive it is: it verifies its own work, builds checks when needed, and keeps iterating on long tasks. it’s available across Claude and the API, with a faster mode that runs around 2.5x the default speed.
🔗 https://www.anthropic.com/news/claude-opus-5
❤4
Jensen Huang made his first post on X a statement of intent: the world needs both open and closed AI models.
he shared a letter signed by NVIDIA, Microsoft, Meta and 22 others arguing that downloadable models give companies more control, lower costs and reduce platform lock-in. the group is urging policymakers to avoid broad restrictions that could weaken competition and push AI development elsewhere.
🔗 https://x.com/JensenHuang/status/2080643682408321103
🔗 https://images.nvidia.com/pdf/Open-Weights-and-American-AI-Leadership.pdf
he shared a letter signed by NVIDIA, Microsoft, Meta and 22 others arguing that downloadable models give companies more control, lower costs and reduce platform lock-in. the group is urging policymakers to avoid broad restrictions that could weaken competition and push AI development elsewhere.
🔗 https://x.com/JensenHuang/status/2080643682408321103
🔗 https://images.nvidia.com/pdf/Open-Weights-and-American-AI-Leadership.pdf
❤5
Media is too big
VIEW IN TELEGRAM
Andon Labs built Drone-Bench to test agents writing code for five physical drone tasks, from reconstructing an office to finding and following a person. each task is isolated with clean upstream artifacts and agents get ten scored attempts.
the strongest models sometimes cleared four tasks, but end-to-end success was still 0% because no run beat the reconstruction baseline. a useful reality check for physical-world agents.
pretty darn cool!
🔗 https://andonlabs.com/evals/drone-bench
the strongest models sometimes cleared four tasks, but end-to-end success was still 0% because no run beat the reconstruction baseline. a useful reality check for physical-world agents.
pretty darn cool!
🔗 https://andonlabs.com/evals/drone-bench
This media is not supported in your browser
VIEW IN TELEGRAM
box is offering persistent Linux VMs for AI agents at $0.036 an hour. each one includes 4 vCPU, 8 GB of RAM, over 50 GB of storage, Docker, SSH, a desktop and snapshots.
the useful part is the fleet model: one account can run up to 1,200 boxes in parallel, while stopped machines keep their files and stop billing. there’s a $20 monthly minimum, which covers around 555 VM hours.
🔗 https://box.ascii.dev/
the useful part is the fleet model: one account can run up to 1,200 boxes in parallel, while stopped machines keep their files and stop billing. there’s a $20 monthly minimum, which covers around 555 VM hours.
🔗 https://box.ascii.dev/