š„Nvidia SpatialClaw is outš„
šFrom Nvidia a novel training-free framework for spatial reasoning that adopts code as the action interface. SpatialClaw lets a VLM-backed agent write Python in a persistent kernel, composing perception modules, inspecting intermediate results, and revising its strategy across steps. Impressive: +11.2 points on 20 benchmarksš
šReview https://t.ly/7JB0x
šPaper https://arxiv.org/pdf/2606.13673
šProject https://spatialclaw.github.io/
šRepo https://github.com/NVlabs/SpatialClaw
šFrom Nvidia a novel training-free framework for spatial reasoning that adopts code as the action interface. SpatialClaw lets a VLM-backed agent write Python in a persistent kernel, composing perception modules, inspecting intermediate results, and revising its strategy across steps. Impressive: +11.2 points on 20 benchmarksš
šReview https://t.ly/7JB0x
šPaper https://arxiv.org/pdf/2606.13673
šProject https://spatialclaw.github.io/
šRepo https://github.com/NVlabs/SpatialClaw
š¤Æ6ā¤2š„2
This media is not supported in your browser
VIEW IN TELEGRAM
šÆWorldwide Semantic FacadešÆ
šA centimeter-accurate / cross-continental facade point clouds, with fine-grained semantic segmentation of architectural elements, and hierarchical facade taxonomy. 2.7B Datasetš
šReview https://t.ly/PpyFD
šPaper https://arxiv.org/pdf/2607.02018
šProject jiangyuanwangyi.github.io/UnderOneFacade_official
šData drive.google.com/drive/folders/1Yzz7PmyeK1qeOtkTFCfkbw7IEHXcMJo8
šA centimeter-accurate / cross-continental facade point clouds, with fine-grained semantic segmentation of architectural elements, and hierarchical facade taxonomy. 2.7B Datasetš
šReview https://t.ly/PpyFD
šPaper https://arxiv.org/pdf/2607.02018
šProject jiangyuanwangyi.github.io/UnderOneFacade_official
šData drive.google.com/drive/folders/1Yzz7PmyeK1qeOtkTFCfkbw7IEHXcMJo8
š¤Æ10ā¤7š„1š1
This media is not supported in your browser
VIEW IN TELEGRAM
šāā¬Spatial-perception native ViTšāā¬
šLingBot-Vision, a vision foundation model pretrained to be spatial-perception native. Better than 7x bigger foundational models. Repo under Apacheš
šReview https://t.ly/9xIso
šPaper https://arxiv.org/pdf/2607.05247
šProject https://technology.robbyant.com/lingbot-vision
šRepo https://github.com/robbyant/lingbot-vision
šLingBot-Vision, a vision foundation model pretrained to be spatial-perception native. Better than 7x bigger foundational models. Repo under Apacheš
šReview https://t.ly/9xIso
šPaper https://arxiv.org/pdf/2607.05247
šProject https://technology.robbyant.com/lingbot-vision
šRepo https://github.com/robbyant/lingbot-vision
š¤Æ7ā¤6š2š1
This media is not supported in your browser
VIEW IN TELEGRAM
šµļøSoccerNet 2026 Resultsšµļø
šThe SoccerNet 2026 Challenges constitute the sixth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision research in sports video understandingš
šReview https://t.ly/sfD4T
šPaper https://lnkd.in/dSBgW_3s
šProject https://lnkd.in/dfdmuvG8
šThe SoccerNet 2026 Challenges constitute the sixth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision research in sports video understandingš
šReview https://t.ly/sfD4T
šPaper https://lnkd.in/dSBgW_3s
šProject https://lnkd.in/dfdmuvG8
š„10ā¤2š2
This media is not supported in your browser
VIEW IN TELEGRAM
š„ZipDepth: Depth on Any Deviceš„
šZipDepth from UniBO is a super-compact monocular depth network by combining an efficient reparameterizable encoder-decoder with large-scale knowledge distillation from a foundation model. Repo under MITš
šReview https://t.ly/qYrLZ
šPaper https://arxiv.org/pdf/2607.08771
šProject https://zipdepth.github.io/
šRepo https://github.com/fabiotosi92/ZipDepth
šZipDepth from UniBO is a super-compact monocular depth network by combining an efficient reparameterizable encoder-decoder with large-scale knowledge distillation from a foundation model. Repo under MITš
šReview https://t.ly/qYrLZ
šPaper https://arxiv.org/pdf/2607.08771
šProject https://zipdepth.github.io/
šRepo https://github.com/fabiotosi92/ZipDepth
ā¤17š„10
This media is not supported in your browser
VIEW IN TELEGRAM
šSAM-MT: Real-Time Multi-Target VOSš
šFudan & Shangai unveil SAM-MT, an efficient interactive multi-target video segmentation framework that maintains near-single-object efficiency (FPS/VRAM) as target count increases, while maintaining robust video segmentation performance. Repo availableš
šReview https://t.ly/Z_4C7
šPaper https://lnkd.in/dvS-iyBD
šProject https://lnkd.in/daQ8na8T
šRepo https://lnkd.in/dgbX2tZv
šFudan & Shangai unveil SAM-MT, an efficient interactive multi-target video segmentation framework that maintains near-single-object efficiency (FPS/VRAM) as target count increases, while maintaining robust video segmentation performance. Repo availableš
šReview https://t.ly/Z_4C7
šPaper https://lnkd.in/dvS-iyBD
šProject https://lnkd.in/daQ8na8T
šRepo https://lnkd.in/dgbX2tZv
ā¤9š„4š1
This media is not supported in your browser
VIEW IN TELEGRAM
šFoundation Global SFMš
šGlob3R is a global SfM-style reconstruction built on 3D foundation models. key idea: explicitly optimize feed-forward geometric predictions. Repo TBAš
šReview https://t.ly/Z_4C7
šPaper https://arxiv.org/pdf/2607.09225
šProject https://junyuandeng.github.io/Glob3r/
šRepo TBA
šGlob3R is a global SfM-style reconstruction built on 3D foundation models. key idea: explicitly optimize feed-forward geometric predictions. Repo TBAš
šReview https://t.ly/Z_4C7
šPaper https://arxiv.org/pdf/2607.09225
šProject https://junyuandeng.github.io/Glob3r/
šRepo TBA
š„10ā¤1š1š1š1
This media is not supported in your browser
VIEW IN TELEGRAM
šREMIND: long-term MOT re-IDš
šREMIND by CVAR-UPM is a novel online tracker designed for long-term multi-object re-ID of generic indoor objects from monocular RGB, requiring neither camera pose nor depth. Repo under MITš
šReview https://t.ly/AkQoI
šPaper https://lnkd.in/dm58mkCv
šProject https://lnkd.in/dZrAZqFe
šRepo https://lnkd.in/dbidrwxU
šREMIND by CVAR-UPM is a novel online tracker designed for long-term multi-object re-ID of generic indoor objects from monocular RGB, requiring neither camera pose nor depth. Repo under MITš
šReview https://t.ly/AkQoI
šPaper https://lnkd.in/dm58mkCv
šProject https://lnkd.in/dZrAZqFe
šRepo https://lnkd.in/dbidrwxU
ā¤6š„4š2š1
This media is not supported in your browser
VIEW IN TELEGRAM
𦧠MonkeyOCRv2 is out! š¦§
šMonkeyOCRv2 is a text-centric visual foundation model that unifies fine-grained text modeling, cross-task representation learning, and cross-lingual generalization in a single encoder. Released for academic research and non-commercial useš
šReview https://t.ly/yicEK
šPaper https://arxiv.org/pdf/2607.11562
šRepo https://github.com/Yuliang-Liu/MonkeyOCRv2
šMonkeyOCRv2 is a text-centric visual foundation model that unifies fine-grained text modeling, cross-task representation learning, and cross-lingual generalization in a single encoder. Released for academic research and non-commercial useš
šReview https://t.ly/yicEK
šPaper https://arxiv.org/pdf/2607.11562
šRepo https://github.com/Yuliang-Liu/MonkeyOCRv2
ā¤13š1š„1
This media is not supported in your browser
VIEW IN TELEGRAM
šFlowWAM: flow->action predictionš
šFlowWAM is a novel dual-stream diffusion framework that adopts optical flow as a unified, video-native action representation. Repo under Apacheš
šReview https://t.ly/FmutT
šPaper https://arxiv.org/abs/2607.13017
šProject https://flow-wam.github.io/
šRepo github.com/YixiangChen515/FlowWAM
šFlowWAM is a novel dual-stream diffusion framework that adopts optical flow as a unified, video-native action representation. Repo under Apacheš
šReview https://t.ly/FmutT
šPaper https://arxiv.org/abs/2607.13017
šProject https://flow-wam.github.io/
šRepo github.com/YixiangChen515/FlowWAM
ā¤8ā”1š1
This media is not supported in your browser
VIEW IN TELEGRAM
šÆSOTA Music-to-Dance GenšÆ
šThe Tongyi Lab unveils Wan-Dancer, a novel stable minute-scale synthesis at 720p/30fps across five dance genres. Impressive results, new SOTA on long clip by a large margin. Repo under Apache 2.0š
šReview https://t.ly/AKY5j
šPaper https://lnkd.in/d_xA7dwb
šProject https://lnkd.in/dzfnw2h4
šRepo https://lnkd.in/d-Zj_cTf
šThe Tongyi Lab unveils Wan-Dancer, a novel stable minute-scale synthesis at 720p/30fps across five dance genres. Impressive results, new SOTA on long clip by a large margin. Repo under Apache 2.0š
šReview https://t.ly/AKY5j
šPaper https://lnkd.in/d_xA7dwb
šProject https://lnkd.in/dzfnw2h4
šRepo https://lnkd.in/d-Zj_cTf
š¤Æ3ā¤2š„1
This media is not supported in your browser
VIEW IN TELEGRAM
šNot a render. Not a concept. This is GENE.01 by Generative Bionics, the Italians coolest scaleup strikes back: in just six months, they turned GENE.01 into a fully functional humanoid platform that can walk, sense and interact.
šFull-body multimodal skin perceives touch, proximity, force and temperature, bringing Physical AI closer to safe and natural collaboration with people.
šMore: https://t.ly/F3I3A
šFull-body multimodal skin perceives touch, proximity, force and temperature, bringing Physical AI closer to safe and natural collaboration with people.
šMore: https://t.ly/F3I3A
š„4ā¤2š¤©2š1
What about more posts about Robotics?
Anonymous Poll
59%
Yes, please ā„ļø
35%
Yes, but only if AI is really relevant within the post
6%
NO, Iām scared about terminator
ā¤3