AI with Papers - Artificial Intelligence & Deep Learning
17.1K subscribers
159 photos
282 videos
14 files
1.46K links
All the AI with papers. Every day fresh updates about #DeepLearning #MachineLearning #LLM & #ComputerVision

Curated by Alessandro Ferrari | https://www.linkedin.com/in/visionarynet/

#AI #chatGPT
Download Telegram
šŸ”„Nvidia SpatialClaw is outšŸ”„

šŸ‘‰From Nvidia a novel training-free framework for spatial reasoning that adopts code as the action interface. SpatialClaw lets a VLM-backed agent write Python in a persistent kernel, composing perception modules, inspecting intermediate results, and revising its strategy across steps. Impressive: +11.2 points on 20 benchmarksšŸ’™

šŸ‘‰Review https://t.ly/7JB0x
šŸ‘‰Paper https://arxiv.org/pdf/2606.13673
šŸ‘‰Project https://spatialclaw.github.io/
šŸ‘‰Repo https://github.com/NVlabs/SpatialClaw
🤯6ā¤2šŸ”„2
This media is not supported in your browser
VIEW IN TELEGRAM
šŸÆWorldwide Semantic FacadešŸÆ

šŸ‘‰A centimeter-accurate / cross-continental facade point clouds, with fine-grained semantic segmentation of architectural elements, and hierarchical facade taxonomy. 2.7B DatasetšŸ’™

šŸ‘‰Review https://t.ly/PpyFD
šŸ‘‰Paper https://arxiv.org/pdf/2607.02018
šŸ‘‰Project jiangyuanwangyi.github.io/UnderOneFacade_official
šŸ‘‰Data drive.google.com/drive/folders/1Yzz7PmyeK1qeOtkTFCfkbw7IEHXcMJo8
🤯10ā¤7šŸ”„1šŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
šŸˆā€ā¬›Spatial-perception native ViTšŸˆā€ā¬›

šŸ‘‰LingBot-Vision, a vision foundation model pretrained to be spatial-perception native. Better than 7x bigger foundational models. Repo under ApachešŸ’™

šŸ‘‰Review https://t.ly/9xIso
šŸ‘‰Paper https://arxiv.org/pdf/2607.05247
šŸ‘‰Project https://technology.robbyant.com/lingbot-vision
šŸ‘‰Repo https://github.com/robbyant/lingbot-vision
🤯7ā¤6šŸ‘2šŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
šŸµļøSoccerNet 2026 ResultsšŸµļø

šŸ‘‰The SoccerNet 2026 Challenges constitute the sixth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision research in sports video understandingšŸ’™

šŸ‘‰Review https://t.ly/sfD4T
šŸ‘‰Paper https://lnkd.in/dSBgW_3s
šŸ‘‰Project https://lnkd.in/dfdmuvG8
šŸ”„10ā¤2šŸ‘2
This media is not supported in your browser
VIEW IN TELEGRAM
šŸ”„ZipDepth: Depth on Any DevicešŸ”„

šŸ‘‰ZipDepth from UniBO is a super-compact monocular depth network by combining an efficient reparameterizable encoder-decoder with large-scale knowledge distillation from a foundation model. Repo under MITšŸ’™

šŸ‘‰Review https://t.ly/qYrLZ
šŸ‘‰Paper https://arxiv.org/pdf/2607.08771
šŸ‘‰Project https://zipdepth.github.io/
šŸ‘‰Repo https://github.com/fabiotosi92/ZipDepth
ā¤17šŸ”„10
This media is not supported in your browser
VIEW IN TELEGRAM
šŸ’‹SAM-MT: Real-Time Multi-Target VOSšŸ’‹

šŸ‘‰Fudan & Shangai unveil SAM-MT, an efficient interactive multi-target video segmentation framework that maintains near-single-object efficiency (FPS/VRAM) as target count increases, while maintaining robust video segmentation performance. Repo availablešŸ’™

šŸ‘‰Review https://t.ly/Z_4C7
šŸ‘‰Paper https://lnkd.in/dvS-iyBD
šŸ‘‰Project https://lnkd.in/daQ8na8T
šŸ‘‰Repo https://lnkd.in/dgbX2tZv
ā¤9šŸ”„4šŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
šŸŒ”Foundation Global SFMšŸŒ”

šŸ‘‰Glob3R is a global SfM-style reconstruction built on 3D foundation models. key idea: explicitly optimize feed-forward geometric predictions. Repo TBAšŸ’™

šŸ‘‰Review https://t.ly/Z_4C7
šŸ‘‰Paper https://arxiv.org/pdf/2607.09225
šŸ‘‰Project https://junyuandeng.github.io/Glob3r/
šŸ‘‰Repo TBA
šŸ”„10ā¤1šŸ‘1šŸ‘1šŸ˜1
This media is not supported in your browser
VIEW IN TELEGRAM
šŸŽ‚REMIND: long-term MOT re-IDšŸŽ‚

šŸ‘‰REMIND by CVAR-UPM is a novel online tracker designed for long-term multi-object re-ID of generic indoor objects from monocular RGB, requiring neither camera pose nor depth. Repo under MITšŸ’™

šŸ‘‰Review https://t.ly/AkQoI
šŸ‘‰Paper https://lnkd.in/dm58mkCv
šŸ‘‰Project https://lnkd.in/dZrAZqFe
šŸ‘‰Repo https://lnkd.in/dbidrwxU
ā¤6šŸ”„4šŸ˜2šŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
🦧 MonkeyOCRv2 is out! 🦧

šŸ‘‰MonkeyOCRv2 is a text-centric visual foundation model that unifies fine-grained text modeling, cross-task representation learning, and cross-lingual generalization in a single encoder. Released for academic research and non-commercial usešŸ’™

šŸ‘‰Review https://t.ly/yicEK
šŸ‘‰Paper https://arxiv.org/pdf/2607.11562
šŸ‘‰Repo https://github.com/Yuliang-Liu/MonkeyOCRv2
ā¤13šŸ‘1šŸ”„1
This media is not supported in your browser
VIEW IN TELEGRAM
🌈FlowWAM: flow->action prediction🌈

šŸ‘‰FlowWAM is a novel dual-stream diffusion framework that adopts optical flow as a unified, video-native action representation. Repo under ApachešŸ’™

šŸ‘‰Review https://t.ly/FmutT
šŸ‘‰Paper https://arxiv.org/abs/2607.13017
šŸ‘‰Project https://flow-wam.github.io/
šŸ‘‰Repo github.com/YixiangChen515/FlowWAM
ā¤8⚔1šŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
šŸÆSOTA Music-to-Dance GenšŸÆ

šŸ‘‰The Tongyi Lab unveils Wan-Dancer, a novel stable minute-scale synthesis at 720p/30fps across five dance genres. Impressive results, new SOTA on long clip by a large margin. Repo under Apache 2.0šŸ’™

šŸ‘‰Review https://t.ly/AKY5j
šŸ‘‰Paper https://lnkd.in/d_xA7dwb
šŸ‘‰Project https://lnkd.in/dzfnw2h4
šŸ‘‰Repo https://lnkd.in/d-Zj_cTf
🤯3ā¤2šŸ”„1
This media is not supported in your browser
VIEW IN TELEGRAM
šŸ‘‰Not a render. Not a concept. This is GENE.01 by Generative Bionics, the Italians coolest scaleup strikes back: in just six months, they turned GENE.01 into a fully functional humanoid platform that can walk, sense and interact.

šŸ‘‰Full-body multimodal skin perceives touch, proximity, force and temperature, bringing Physical AI closer to safe and natural collaboration with people.

šŸ‘‰More: https://t.ly/F3I3A
šŸ”„4ā¤2🤩2šŸ‘1