AI with Papers - Artificial Intelligence & Deep Learning
17.2K subscribers
163 photos
287 videos
14 files
1.5K links
All the AI with papers. Every day fresh updates about #DeepLearning #MachineLearning #LLM & #ComputerVision

Curated by Alessandro Ferrari | https://www.linkedin.com/in/visionarynet/

#AI #chatGPT
Download Telegram
This media is not supported in your browser
VIEW IN TELEGRAM
šŸ”„ RelateAnything is gold! šŸ”„

šŸ‘‰RelateAnything is a 53M-parameter relation model that takes an image and a set of regions from any source and returns scored relations over a predicate vocabulary supplied at inference as a list of strings. Impressive results. Repo under Apache 2.0šŸ’™

šŸ‘‰Review https://lnkd.in/p/etAcdFM3
šŸ‘‰Paper https://arxiv.org/pdf/2609.12552
šŸ‘‰Repo https://github.com/Maelic/RelateAnything
šŸ‘‰Project https://maelic.github.io/RelateAnythingProject/
ā¤12šŸ”„4šŸ‘3šŸ‘1🤯1
This media is not supported in your browser
VIEW IN TELEGRAM
šŸ‘‹ EventEgoHands++ is out! šŸ‘‹

šŸ‘‰EventEgoHands++ is a novel framework for event-based 3D hand mesh reconstruction from an egocentric viewpoint. 1M+ samples dataset! Code/Data releasedšŸ’™

šŸ‘‰Review https://lnkd.in/p/eTbPvXbW
šŸ‘‰Paper https://arxiv.org/pdf/2609.17189
šŸ‘‰Repo https://github.com/ryhara/EventEgoHandsV2
šŸ‘‰Project https://ryhara.github.io/EventEgoHandsV2/
ā¤3šŸ”„2šŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
šŸ’¦SOTA Splashing LiquidsšŸ’¦

šŸ‘‰SplashSplat reconstructs splashing liquids from real multi-view vide. Impose physical structure only where the observations can constrain it. Impressive results, SOTA. Code TBR under MITšŸ’™

šŸ‘‰Review https://lnkd.in/p/ejMTHcp7
šŸ‘‰Paper https://arxiv.org/pdf/2609.20818
šŸ‘‰Project niko-creater.github.io/splashsplat-web/
šŸ‘‰Repo https://github.com/Niko-creater/Splashsplat
šŸ‘4ā¤2šŸ”„2šŸ‘1
+++ Breaking +++
ā¤3šŸ‘2😢2🤣1
This media is not supported in your browser
VIEW IN TELEGRAM
šŸ”„Agentic Image-to-ScenešŸ”„

šŸ‘‰HARMONY by UPenn is a hierarchical chain-of-thought framework that leverages both agentic reasoning and visual geometry foundation. Impressive 3D scenes. Repo TBAšŸ’™

šŸ‘‰Review https://lnkd.in/p/ep2hmRSp
šŸ‘‰Paper https://arxiv.org/pdf/2609.26793
šŸ‘‰Project https://cwchenwang.github.io/harmony/
šŸ‘‰Data https://huggingface.co/datasets/ShufanSun/harmony
šŸ”„8ā¤2šŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
šŸæPanoSeg3R: SOTA 3D SegmentationšŸæ

šŸ‘‰PanoSeg3R is a novel feed-forward framework for 3D panoramic semantic segmentation. New SOTA. Code comingšŸ’™

šŸ‘‰Review https://lnkd.in/p/eKCKWv3g
šŸ‘‰Paper https://arxiv.org/pdf/2609.22687
šŸ‘‰Project https://harryyoon777.github.io/PanoSeg3R/#
šŸ‘‰Repo TBA
ā¤5šŸ‘1šŸ”„1šŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
🩻Universal X-ray Segmentation🩻

šŸ‘‰FleXray: universal anatomical segmentation across the entire body in clinical X-rays. Built on a scalable, physics-based generative X-ray data engine. Repo under MITšŸ’™

šŸ‘‰Review https://lnkd.in/p/e9MUk_eq
šŸ‘‰Paper https://arxiv.org/pdf/2609.26756
šŸ‘‰Project https://flexray.csail.mit.edu/
šŸ‘‰Repo https://github.com/VictorButoi/FleXray
šŸ‘5ā¤3šŸ”„3šŸ‘2🤯1
This media is not supported in your browser
VIEW IN TELEGRAM
🦓3D Foundational Radiology🦓

šŸ‘‰nnFoundation: 3D radiological foundation models designed for transferable representation learning across heterogeneous tasks/datasets. Models releasedšŸ’™

šŸ‘‰Review https://lnkd.in/p/eNajRGBi
šŸ‘‰Paper https://arxiv.org/pdf/2609.26924
šŸ‘‰Models https://huggingface.co/collections/MIC-DKFZ/nnfoundation
ā¤9šŸ‘2šŸ‘2šŸ”„1🤩1
This media is not supported in your browser
VIEW IN TELEGRAM
šŸ”„TrackEverything is outšŸ”„

šŸ‘‰TrackEverything is the first 3D point tracker capable of tracking all visible points across long horizons (1000+ frames). Repo announcedšŸ’™

šŸ‘‰Review https://lnkd.in/p/eCPJ6h2B
šŸ‘‰Paper https://arxiv.org/pdf/2609.30222
šŸ‘‰Project https://trackeverything.github.io/
šŸ‘‰Repo https://github.com/ayushjain1144/trackeverything
ā¤6šŸ”„6šŸ‘2šŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
šŸ”„Ego-Exo4D Human DatasetšŸ”„

šŸ‘‰Form the University of Austin, Ego-Exo4D-HM: large-scale dataset of 4D human motion reconstructions for Ego-Exo4D’s captures + reconstruction pipeline. Code, dataset, and docs šŸ’™

šŸ‘‰Review https://lnkd.in/p/eVFt9jPr
šŸ‘‰Paper https://lnkd.in/eWj4cD7T
šŸ‘‰Project https://lnkd.in/euPqVNxV
1ā¤5šŸ”„3šŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
šŸ”„šŸ”„ 70,000+ šŸ”„šŸ”„

šŸ‘‰ Crazy how a boring science project (no kittens, no rants, no personal dramas) can reach for 70,000+ followers. Speechless.

Love u šŸ’›

šŸ‘‰ https://lnkd.in/p/eD6Xxdxi
ā¤26šŸ¾8šŸ”„3⚔2šŸ‘2šŸ‘2🤯1
šŸ”„The Computer Vision ultimate collectionšŸ”„

šŸ‘‰Stan Birchfield (#Nvidia) just dropped this on arXiv. From classical image processing and 3D geometry to CNNs, Transformers, foundation models, and neural rendering. What makes this book damn good is the combination of clear explanations and working Python. A gift.

šŸ‘‰Review https://lnkd.in/p/ejwm_DVn
šŸ‘‰Book https://lnkd.in/eTrEvmd9
šŸ‘‰Code https://lnkd.in/eakj9VZU
ā¤26šŸ”„8šŸ‘2šŸ’©1šŸ˜1
This media is not supported in your browser
VIEW IN TELEGRAM
šŸ‡Physically Plausible 3D MotionšŸ‡

šŸ‘‰Physically plausible motion recovery: given a monocular video, FlowHMR recovers global 3D human motion that a physics-based controller can successfully track in simulation. RepošŸ’™

šŸ‘‰Review https://lnkd.in/p/d77fzUtR
šŸ‘‰Paper https://arxiv.org/pdf/2610.03691
šŸ‘‰Project https://flowhmr.github.io/
šŸ‘‰Repo https://github.com/flowhmr/flowhmr
ā¤6šŸ”„4šŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
🟨 Rome from ONE pic 🟄

šŸ‘‰Detailed scene meshes from one photograph: the method completes geometry beyond the observed view and supports indoor, outdoor, and large-scale scenes. Repo announcedšŸ’™

šŸ‘‰Review https://lnkd.in/p/eMdWBEGn
šŸ‘‰Paper https://arxiv.org/pdf/2610.08790
šŸ‘‰Project https://build-rome.github.io/
šŸ‘‰Repo TBA
ā¤7🤣3