AI with Papers - Artificial Intelligence & Deep Learning
17.2K subscribers
163 photos
287 videos
14 files
1.5K links
All the AI with papers. Every day fresh updates about #DeepLearning #MachineLearning #LLM & #ComputerVision

Curated by Alessandro Ferrari | https://www.linkedin.com/in/visionarynet/

#AI #chatGPT
Download Telegram
This media is not supported in your browser
VIEW IN TELEGRAM
🦺Efficient/Scalable Video Pretraining🦺

πŸ‘‰LeVJEPA1 (Yann Lecun) is the first video encoder trained under LeJEPA’s collapse-free objective, and evaluate it under frozen probing against video and image pretraining baselines retrained on identical data, in both epoch-matched and FLOP-matched regimes. Repo under MITπŸ’™

πŸ‘‰Review https://lnkd.in/p/eJQAm3AN
πŸ‘‰Paper https://lnkd.in/eCzzTiNH
πŸ‘‰Project https://levjepa.github.io/
πŸ‘‰Repo https://lnkd.in/etiF5CDj
❀8πŸ”₯5πŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ”₯ RelateAnything is gold! πŸ”₯

πŸ‘‰RelateAnything is a 53M-parameter relation model that takes an image and a set of regions from any source and returns scored relations over a predicate vocabulary supplied at inference as a list of strings. Impressive results. Repo under Apache 2.0πŸ’™

πŸ‘‰Review https://lnkd.in/p/etAcdFM3
πŸ‘‰Paper https://arxiv.org/pdf/2609.12552
πŸ‘‰Repo https://github.com/Maelic/RelateAnything
πŸ‘‰Project https://maelic.github.io/RelateAnythingProject/
❀12πŸ”₯4πŸ‘3πŸ‘1🀯1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ‘‹ EventEgoHands++ is out! πŸ‘‹

πŸ‘‰EventEgoHands++ is a novel framework for event-based 3D hand mesh reconstruction from an egocentric viewpoint. 1M+ samples dataset! Code/Data releasedπŸ’™

πŸ‘‰Review https://lnkd.in/p/eTbPvXbW
πŸ‘‰Paper https://arxiv.org/pdf/2609.17189
πŸ‘‰Repo https://github.com/ryhara/EventEgoHandsV2
πŸ‘‰Project https://ryhara.github.io/EventEgoHandsV2/
❀3πŸ”₯2πŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ’¦SOTA Splashing LiquidsπŸ’¦

πŸ‘‰SplashSplat reconstructs splashing liquids from real multi-view vide. Impose physical structure only where the observations can constrain it. Impressive results, SOTA. Code TBR under MITπŸ’™

πŸ‘‰Review https://lnkd.in/p/ejMTHcp7
πŸ‘‰Paper https://arxiv.org/pdf/2609.20818
πŸ‘‰Project niko-creater.github.io/splashsplat-web/
πŸ‘‰Repo https://github.com/Niko-creater/Splashsplat
πŸ‘4❀2πŸ”₯2πŸ‘1
+++ Breaking +++
❀3πŸ‘2😒2🀣1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ”₯Agentic Image-to-SceneπŸ”₯

πŸ‘‰HARMONY by UPenn is a hierarchical chain-of-thought framework that leverages both agentic reasoning and visual geometry foundation. Impressive 3D scenes. Repo TBAπŸ’™

πŸ‘‰Review https://lnkd.in/p/ep2hmRSp
πŸ‘‰Paper https://arxiv.org/pdf/2609.26793
πŸ‘‰Project https://cwchenwang.github.io/harmony/
πŸ‘‰Data https://huggingface.co/datasets/ShufanSun/harmony
πŸ”₯8❀2πŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
🍿PanoSeg3R: SOTA 3D Segmentation🍿

πŸ‘‰PanoSeg3R is a novel feed-forward framework for 3D panoramic semantic segmentation. New SOTA. Code comingπŸ’™

πŸ‘‰Review https://lnkd.in/p/eKCKWv3g
πŸ‘‰Paper https://arxiv.org/pdf/2609.22687
πŸ‘‰Project https://harryyoon777.github.io/PanoSeg3R/#
πŸ‘‰Repo TBA
❀5πŸ‘1πŸ”₯1πŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
🩻Universal X-ray Segmentation🩻

πŸ‘‰FleXray: universal anatomical segmentation across the entire body in clinical X-rays. Built on a scalable, physics-based generative X-ray data engine. Repo under MITπŸ’™

πŸ‘‰Review https://lnkd.in/p/e9MUk_eq
πŸ‘‰Paper https://arxiv.org/pdf/2609.26756
πŸ‘‰Project https://flexray.csail.mit.edu/
πŸ‘‰Repo https://github.com/VictorButoi/FleXray
πŸ‘5❀3πŸ”₯3πŸ‘2🀯1
This media is not supported in your browser
VIEW IN TELEGRAM
🦴3D Foundational Radiology🦴

πŸ‘‰nnFoundation: 3D radiological foundation models designed for transferable representation learning across heterogeneous tasks/datasets. Models releasedπŸ’™

πŸ‘‰Review https://lnkd.in/p/eNajRGBi
πŸ‘‰Paper https://arxiv.org/pdf/2609.26924
πŸ‘‰Models https://huggingface.co/collections/MIC-DKFZ/nnfoundation
❀9πŸ‘2πŸ‘2πŸ”₯1🀩1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ”₯TrackEverything is outπŸ”₯

πŸ‘‰TrackEverything is the first 3D point tracker capable of tracking all visible points across long horizons (1000+ frames). Repo announcedπŸ’™

πŸ‘‰Review https://lnkd.in/p/eCPJ6h2B
πŸ‘‰Paper https://arxiv.org/pdf/2609.30222
πŸ‘‰Project https://trackeverything.github.io/
πŸ‘‰Repo https://github.com/ayushjain1144/trackeverything
❀6πŸ”₯6πŸ‘2πŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ”₯Ego-Exo4D Human DatasetπŸ”₯

πŸ‘‰Form the University of Austin, Ego-Exo4D-HM: large-scale dataset of 4D human motion reconstructions for Ego-Exo4D’s captures + reconstruction pipeline. Code, dataset, and docs πŸ’™

πŸ‘‰Review https://lnkd.in/p/eVFt9jPr
πŸ‘‰Paper https://lnkd.in/eWj4cD7T
πŸ‘‰Project https://lnkd.in/euPqVNxV
1❀5πŸ”₯3πŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ”₯πŸ”₯ 70,000+ πŸ”₯πŸ”₯

πŸ‘‰ Crazy how a boring science project (no kittens, no rants, no personal dramas) can reach for 70,000+ followers. Speechless.

Love u πŸ’›

πŸ‘‰ https://lnkd.in/p/eD6Xxdxi
❀26🍾8πŸ”₯3⚑2πŸ‘2πŸ‘2🀯1
πŸ”₯The Computer Vision ultimate collectionπŸ”₯

πŸ‘‰Stan Birchfield (#Nvidia) just dropped this on arXiv. From classical image processing and 3D geometry to CNNs, Transformers, foundation models, and neural rendering. What makes this book damn good is the combination of clear explanations and working Python. A gift.

πŸ‘‰Review https://lnkd.in/p/ejwm_DVn
πŸ‘‰Book https://lnkd.in/eTrEvmd9
πŸ‘‰Code https://lnkd.in/eakj9VZU
❀26πŸ”₯8πŸ‘2πŸ’©1😍1
This media is not supported in your browser
VIEW IN TELEGRAM
πŸ‡Physically Plausible 3D MotionπŸ‡

πŸ‘‰Physically plausible motion recovery: given a monocular video, FlowHMR recovers global 3D human motion that a physics-based controller can successfully track in simulation. RepoπŸ’™

πŸ‘‰Review https://lnkd.in/p/d77fzUtR
πŸ‘‰Paper https://arxiv.org/pdf/2610.03691
πŸ‘‰Project https://flowhmr.github.io/
πŸ‘‰Repo https://github.com/flowhmr/flowhmr
❀6πŸ”₯4πŸ‘1
This media is not supported in your browser
VIEW IN TELEGRAM
🟨 Rome from ONE pic πŸŸ₯

πŸ‘‰Detailed scene meshes from one photograph: the method completes geometry beyond the observed view and supports indoor, outdoor, and large-scale scenes. Repo announcedπŸ’™

πŸ‘‰Review https://lnkd.in/p/eMdWBEGn
πŸ‘‰Paper https://arxiv.org/pdf/2610.08790
πŸ‘‰Project https://build-rome.github.io/
πŸ‘‰Repo TBA
❀7🀣3